🔧 AI Nachrichten Major AI platforms go down in unprecedented simultaneous outage(03.09.2026 um 17:34 Uhr)
🔧 AI Nachrichten ChatGPT, Claude, and Grok Down? Users Report Widespread Outages(03.09.2026 um 19:14 Uhr)
🔧 AI Nachrichten OpenAI Launches GPT-6 Astra, Says We May Have Entered the AGI Era(03.09.2026 um 22:08 Uhr)
🔧 AI Nachrichten Claude Comes to CarPlay as Fifth Major AI Chatbot App(05.09.2026 um 05:31 Uhr)
🔧 AI Nachrichten OpenAI’s GPT-6 Astra Is AGI, Says NVIDIA CEO Jensen Huang(07.09.2026 um 06:31 Uhr)
🔧 AI Nachrichten Blame AI companies for Mac mini and Mac Studio shortage(31.08.2026 um 10:32 Uhr)
🔧 AI Nachrichten Major AI platforms go down in unprecedented simultaneous outage(03.09.2026 um 17:34 Uhr)
🔧 AI Nachrichten ChatGPT, Claude, and Grok Down? Users Report Widespread Outages(03.09.2026 um 19:14 Uhr)
🔧 AI Nachrichten OpenAI Launches GPT-6 Astra, Says We May Have Entered the AGI Era(03.09.2026 um 22:08 Uhr)
🔧 AI Nachrichten Claude Comes to CarPlay as Fifth Major AI Chatbot App(05.09.2026 um 05:31 Uhr)
🔧 AI Nachrichten OpenAI’s GPT-6 Astra Is AGI, Says NVIDIA CEO Jensen Huang(07.09.2026 um 06:31 Uhr)
🔧 AI Nachrichten Blame AI companies for Mac mini and Mac Studio shortage(31.08.2026 um 10:32 Uhr)

🔧 Programmierung 🕛 kürzlich 4 Min Lesezeit
0

I Ditched Vector Search for My Coding Agent's Memory. FTS5 Won.

↗ Quelle (dev.to)
🗣️ Stimme:
📑 Inhaltsübersicht

Every "give your agent memory" tutorial I've read reaches for the same stack: chunk your docs, embed them, throw the vectors in a database, do cosine similarity at query time. So when I needed my coding agent to search through indexed tool output, git logs, and fetched docs without dumping raw text into the model's context window, I assumed I'd be standing up a vector store too.



I didn't. I used SQLite's FTS5 full-text search instead, and for this specific job it's not a compromise — it's the better tool.






What the problem actually was



The tool I built (context-mode, for routing large command output and API responses out of the model's context) needs to answer queries like:




  • "failing tests"

  • "HTTP 500 errors"

  • "async route handlers"



against arbitrary shell output, JSON responses, and fetched web pages — indexed once, searched however many times a session needs. The naive version just dumps everything into context and lets the model read it. That works until the output is 50KB of test logs and you've burned half your context window on a summary you needed three lines of.






Why vectors are the wrong default here, not just an alternative



Vector search is built to answer "what's semantically similar to this." That's the right tool when you're searching prose — support tickets, documentation, chat transcripts — where the same idea gets expressed in different words and you need "how do I reset my password" to match a doc titled "Account Recovery Steps."



Coding-agent queries mostly aren't that. "HTTP 500 errors" isn't a fuzzy semantic concept I want approximated — it's closer to a literal grep with better ranking. The content being searched is also structured and keyword-dense: stack traces, log lines, JSON keys, error codes. Embedding a stack trace and comparing cosine similarity throws away the thing that actually matters (the literal exception name, the literal line number) in favor of a vector representation that's better at "these two paragraphs are about similar topics" than "this line contains the string ECONNREFUSED."



FTS5 is built for exactly this: tokenized, indexed, ranked full-text search over exact and near-exact term matches, with BM25-style relevance scoring out of the box.






What it actually looks like



No embedding model, no vector database, no network round-trip to compute embeddings. It's stdlib:




CODE
import sqlite3

conn = sqlite3.connect("index.db")
conn.execute("""
CREATE VIRTUAL TABLE IF NOT EXISTS docs
USING fts5(source, content)
""")

def index(source: str, content: str):
conn.execute("INSERT INTO docs (source, content) VALUES (?, ?)", (source, content))
conn.commit()

def search(query: str, limit: int = 5):
rows = conn.execute("""
SELECT source, snippet(docs, 1,
'[', ']', '...', 20), rank
FROM docs WHERE docs MATCH ? ORDER BY rank LIMIT ?
""", (query, limit)).fetchall()
return rows






That's the whole engine. snippet() gives you highlighted context around the match for free. rank gives you BM25 ordering for free. Querying "HTTP 500 errors" against a batch of indexed test output returns the actual lines containing 500 and error, ranked by term frequency and rarity — not the semantically-nearest paragraph, the actually-relevant one.






Where this would fall over — and why it doesn't here



FTS5 is a bad choice if your queries genuinely need semantic matching: "find the doc about resetting my password" needs to match "Account Recovery," and no amount of tokenization gets you there without embeddings. If I were building search over a knowledge base of prose documentation with inconsistent terminology, I'd reach for vectors, possibly hybrid (BM25 for recall, vectors for semantic re-ranking).



But an agent's own tool output, error logs, and fetched API responses are dense with the literal terms you're going to search for, because you (or the agent) wrote the query with those terms in mind. "Failing tests" as a query is going to co-occur with FAIL, AssertionError, test names — words that are actually in the log. The semantic gap that justifies embeddings mostly doesn't exist in this domain.






The generalizable lesson



"Add semantic search" has become a reflex the same way "add a cache" or "add a queue" is — reached for because it's the default answer to "how do I search this," not because the problem demands it. Vector infra costs you an embedding model, a vector database or extension, and a slower indexing step, in exchange for a capability — semantic similarity — that keyword-dense, structured content usually doesn't need.



Before reaching for embeddings on your next "agent needs to search X" problem, ask what the query and the content actually look like. If both are keyword-dense and structurally similar (logs, code, JSON, stack traces), full-text search with BM25 ranking will outperform vectors on relevance and cost you a fraction of the infrastructure. Save the vector database for the day your content is actually prose with vocabulary mismatch — most agent tooling isn't there yet.

Vollständiger Original-Bericht
Ausführliche Details, Code-Beispiele & Hersteller-Stellungnahme auf dev.to.
↗ Original-Artikel auf dev.to lesen
Wie bewertest du diesen Beitrag?
1 Klick Feedback
Teilen mit Netzwerk & Team:

Community-Analysen & Experten-Meinungen 0

Verfasse deine eigene Analyse, teile Workarounds oder diskutiere diesen Vorfall im Blog.
Noch keine Community-Analyse verfasst. Markiere einen Textabschnitt oder klicke oben auf Eigene Analyse verfassen“!
Community Pulse: Relevanz-Einschätzung
1 Klick Experten-Votum
🔴 Akute Relevanz 0%
🟡 In Evaluierung 0%
🟢 Keine Auswirkung 0%
Spannende Innovation 0%
Verwandte Story-Cluster & Quellen (Vektor-KI)
Port 8095 Engine
3 Quellen
GPT-6 Astra Release Today? OpenAI’s Next Major AI Model Is Almost Here
1 Quelle
Apple accuses OpenAI of destroying evidence as trade-secrets fight intensifies
1 Quelle
Major AI platforms go down in unprecedented simultaneous outage
Ähnliche Beiträge
🔍 Verwandte News

Auch interessante Nachrichten I Ditched Vector Search for My Coding Agent's Memory. FTS5 Won.

Thematisch verwandte Begriffe: Ditched, Vector, Search, Coding · 6 Treffer

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...