📰 IT NachrichtenToday’s NYT Mini Crossword Answers for Saturay, Sept. 12(12.09.2026 um 07:43 Uhr)
🔧 AI Nachrichten Etzioni on AI: What kids tell chatbots, but not you(04.09.2026 um 16:05 Uhr)
🔧 AI Nachrichten OpenAI Wants to Know if an AI Industry Slowdown Would Even Be Legal(11.09.2026 um 01:28 Uhr)
🔧 AI Nachrichten OpenAI puts Pro subscriptions on hold due to Astra demand(10.09.2026 um 22:59 Uhr)
🔧 AI Nachrichten OpenAI’s feud with mathematicians is only escalating(11.09.2026 um 22:57 Uhr)
📰 IT NachrichtenToday’s NYT Mini Crossword Answers for Saturay, Sept. 12(12.09.2026 um 07:43 Uhr)
🔧 AI Nachrichten Etzioni on AI: What kids tell chatbots, but not you(04.09.2026 um 16:05 Uhr)
🔧 AI Nachrichten OpenAI Wants to Know if an AI Industry Slowdown Would Even Be Legal(11.09.2026 um 01:28 Uhr)
🔧 AI Nachrichten OpenAI puts Pro subscriptions on hold due to Astra demand(10.09.2026 um 22:59 Uhr)
🔧 AI Nachrichten OpenAI’s feud with mathematicians is only escalating(11.09.2026 um 22:57 Uhr)

🔧 Programmierung 🕛 vor 1 Monat 3 Min Lesezeit
0

Part 1-3: From Raw Health Text to a Working Retrieval Pipeline

↗ Quelle (dev.to)
🗣️ Stimme:
📑 Inhaltsübersicht

I'm a first-year undergraduate student in Artificial Intelligence Engineering, and this summer I'm building a small RAG (Retrieval-Augmented Generation) project from scratch to learn how these systems actually work under the hood — not just by calling an API, but by building the pipeline piece by piece. This is a log of the first three parts of that journey.






Part 1: Understanding the Building Blocks



Before writing any real project code, I spent time understanding the core ideas behind RAG:





  • Calling an LLM API: sending a prompt programmatically and getting a response back. I used the Gemini API for this. llm_test.py


  • Embeddings: the idea that text can be converted into numerical vectors, where semantically similar sentences end up close to each other in vector space. I tested this with a handful of sentences using sentence-transformers and computed cosine similarity between them — it was satisfying to see related sentences actually cluster together numerically. embedding_test.py


  • Why not just dump everything into the LLM?: context windows are limited, it's expensive, and irrelevant information can actually hurt answer quality rather than help it.


  • Vector databases: conceptually, tools like Chroma or FAISS solve the problem of searching through large numbers of embeddings quickly to find the most relevant ones.
    No heavy coding in Part 1 — mostly small test scripts to confirm I understood each piece before combining them.






Part 2: Collecting and Preparing Real Data



With the concepts in place, Part 2 was about getting real data ready for retrieval:





  • Data collection: I gathered short health topic descriptions (e.g. diabetes, epilepsy) and saved them as structured JSON records, each with a source and text field. Keeping the source attached to each record matters — it means the system can eventually point back to where an answer came from. data/health_data.json


  • Chunking: I wrote a script to split each record's text into smaller, paragraph-level pieces while keeping the original source attached to every chunk. Since I intentionally kept the raw text short for this first test run, most entries ended up as a single chunk each — which is fine for now, since the goal was to validate the pipeline, not build a large dataset yet.


  • Why chunk at all?: embedding models represent short, focused pieces of text more accurately than long documents. Splitting text into meaningful chunks is what makes retrieval actually useful later.






Part 3: Setting Up the Vector Store and First Retrieval



With chunked, source-tagged data ready, Part 3 turned that data into something actually searchable:





  • Vector store setup: installed and configured Chroma, then embedded every chunk from Part 2 and stored it alongside its text and source metadata. embed_and_store.py


  • First retrieval test: wrote a query script, asked a sample question, and retrieved the most similar chunks from the vector store. The results were reasonable given how small and short the dataset still is — a good early sign that the pipeline itself works end to end. retrieval_test.py


  • A note on limitations: since the dataset is intentionally tiny at this stage, retrieval quality is limited. That's expected, and it'll improve as the dataset grows in later parts.






What's Next



At this point I have a full (if small-scale) pipeline: raw text → structured data → chunks → embeddings → retrieval. The next step is connecting retrieval to actual answer generation — feeding retrieved chunks into the LLM as context so it can answer health questions grounded in real source material.



👉 Code for this project: health-rag-assistant






Follow along as I build this project part by part — code on GitHub, progress here on Dev.to.

Vollständiger Original-Bericht
Ausführliche Details, Code-Beispiele & Hersteller-Stellungnahme auf dev.to.
↗ Original-Artikel auf dev.to lesen
Wie bewertest du diesen Beitrag?
1 Klick Feedback
Teilen mit Netzwerk & Team:

Community-Analysen & Experten-Meinungen 0

Verfasse deine eigene Analyse, teile Workarounds oder diskutiere diesen Vorfall im Blog.
Noch keine Community-Analyse verfasst. Markiere einen Textabschnitt oder klicke oben auf Eigene Analyse verfassen“!
Community Pulse: Relevanz-Einschätzung
1 Klick Experten-Votum
🔴 Akute Relevanz 0%
🟡 In Evaluierung 0%
🟢 Keine Auswirkung 0%
Spannende Innovation 0%
Verwandte Story-Cluster & Quellen (Vektor-KI)
Port 8095 Engine
2 Quellen
Seattle Times sues Microsoft and OpenAI, alleging they trained their AI on its journalism
1 Quelle
Today’s NYT Mini Crossword Answers for Saturay, Sept. 12
1 Quelle
Etzioni on AI: What kids tell chatbots, but not you
Ähnliche Beiträge
🔍 Verwandte News

Auch interessante Nachrichten Part 1-3: From Raw Health Text to a Working Retrieval Pipeline

Thematisch verwandte Begriffe: Part, From, Health, Text · 6 Treffer

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...