✦ Case study · AI tooling

Semantic Code Search for AI Agents

AI coding agents read like someone searching a warehouse with a floodlight — everything lit, nothing found. This MCP server hands them a lantern: retrieve exactly the lines that matter, from nine enterprise repos, fully on-prem.

Context

AI coding assistants working across a large enterprise codebase — nine repositories of C#, Java, and TypeScript — kept doing the expensive thing: reading entire files to find one method, scanning directories to locate a class, burning tokens on code that never mattered to the task. On big enterprise files that meant thousands of wasted tokens per question, slower answers, and context windows full of noise.

The fix wasn't a bigger context window. It was retrieval: give the agent tools to ask for exactly what it needs, and make the cheap path the default.

Constraints

Architecture

VS Code opens · file saves Indexer SHA cache · 60-line chunks Ollama local embeddings · 768-d Qdrant vectors · cosine AI Agent MCP client MCP Server 5 tiered tools Filesystem structure · line ranges changed files chunks upsert "payment flow?" cheap reads semantic search paths + line ranges

light = code knowledge. amber: what the agent asks · blue: how the index stays fresh

The tiered-cost toolbox

The design insight: most questions don't need embeddings. The server exposes five tools ordered by cost, and the agent's instructions make the cheap path the default — semantic search is the escalation, not the reflex.

ToolCostUse
list_reposfreeWhat repos exist, with source file counts
keyword_searchfreeExact class / method / error-code lookup — plain text scan, no vectors
get_file_structurefreeClasses, methods, interfaces with line numbers — regex parsers per language, no file content returned
get_chunkfreeRead one line range — the method, not the file
search_codeNatural-language search — "JWT authentication filter" → ranked chunks across all repos

A typical flow costs almost nothing: keyword_search("PaymentController") finds the file, get_file_structure maps it, get_chunk(45, 90) reads the one method that matters. The agent read ~45 lines instead of a 2,000-line file.

How the index stays fresh

Why it's built this way

The hard requirement was trust with enterprise source code — the same constraint as every AI system worth shipping inside a company. A cloud embedding API would have been three lines of code and an immediate compliance violation. Running Ollama for embeddings and Qdrant for vectors, both local, means the index is as private as the code itself.

The second insight is about agent behavior: tools alone don't change habits. The coding agent's mode instructions were rewritten so retrieval is mandatory before any file read — search first, read line ranges second, whole files never. That's what actually moved the token needle: the workflow, enforced, not just available.

Stack

Python · MCP (Model Context Protocol) · Qdrant · Ollama · nomic-embed-text · watchdog · asyncio · Docker

← All projects Discuss a similar problem