Semantic Code Search for AI Agents
AI coding agents read like someone searching a warehouse with a floodlight — everything lit, nothing found. This MCP server hands them a lantern: retrieve exactly the lines that matter, from nine enterprise repos, fully on-prem.
Context
AI coding assistants working across a large enterprise codebase — nine repositories of C#, Java, and TypeScript — kept doing the expensive thing: reading entire files to find one method, scanning directories to locate a class, burning tokens on code that never mattered to the task. On big enterprise files that meant thousands of wasted tokens per question, slower answers, and context windows full of noise.
The fix wasn't a bigger context window. It was retrieval: give the agent tools to ask for exactly what it needs, and make the cheap path the default.
Constraints
- Code never leaves the machine — enterprise source can't go to a cloud embedding API. Embeddings run on local Ollama, vectors live in a local Qdrant container. Zero external calls.
- Zero per-repo configuration — developers shouldn't maintain index configs. The indexer watches VS Code itself: open a project, it gets indexed. Save a file, that file re-indexes. Nothing else to remember.
- Commodity hardware — no GPU cluster. The whole stack runs beside the IDE on a workstation with an entry-level GPU.
- Incremental, always — full re-indexing on every change would make the index perpetually stale. SHA-1 hash caching means only changed files are ever re-embedded.
Architecture
light = code knowledge. amber: what the agent asks · blue: how the index stays fresh
The tiered-cost toolbox
The design insight: most questions don't need embeddings. The server exposes five tools ordered by cost, and the agent's instructions make the cheap path the default — semantic search is the escalation, not the reflex.
| Tool | Cost | Use |
|---|---|---|
| list_repos | free | What repos exist, with source file counts |
| keyword_search | free | Exact class / method / error-code lookup — plain text scan, no vectors |
| get_file_structure | free | Classes, methods, interfaces with line numbers — regex parsers per language, no file content returned |
| get_chunk | free | Read one line range — the method, not the file |
| search_code | embedding | Natural-language search — "JWT authentication filter" → ranked chunks across all repos |
A typical flow costs almost nothing: keyword_search("PaymentController") finds the file, get_file_structure maps it, get_chunk(45, 90) reads the one method that matters. The agent read ~45 lines instead of a 2,000-line file.
How the index stays fresh
- It watches the IDE, not a config file. The indexer reads VS Code's workspace state — open a project and indexing starts, no registration step. Only the projects actually being worked on get indexed.
- Save a file, that file re-indexes. A filesystem watcher catches every save and delete; changed files are re-chunked, re-embedded, and upserted within seconds.
- Hashes gate everything. Every file's SHA-1 is cached; unchanged files are never re-embedded. A repo re-open with nothing changed costs zero embedding calls.
- Chunks overlap. 60-line chunks with 10-line overlap, so a method spanning a boundary is still findable from either side. Payload indexes on repo and file path make filtered queries fast.
- Concurrency tuned to the metal. Ten parallel embedding calls saturate a small GPU without falling over; vector upserts batch in the hundreds to cut round-trips.
Why it's built this way
The hard requirement was trust with enterprise source code — the same constraint as every AI system worth shipping inside a company. A cloud embedding API would have been three lines of code and an immediate compliance violation. Running Ollama for embeddings and Qdrant for vectors, both local, means the index is as private as the code itself.
The second insight is about agent behavior: tools alone don't change habits. The coding agent's mode instructions were rewritten so retrieval is mandatory before any file read — search first, read line ranges second, whole files never. That's what actually moved the token needle: the workflow, enforced, not just available.