🔧 ProgrammierungGitHub Release: dependabot/dependabot-core v0.395.0 (07.09.2026)(07.09.2026 um 15:21 Uhr)
🔧 ProgrammierungGitHub Release: dependabot/dependabot-core v0.396.0 (14.09.2026)(14.09.2026 um 19:05 Uhr)
🔧 ProgrammierungGitHub Release: langwatch/scenario vpython/v1.4.0 (02.09.2026)(02.09.2026 um 04:47 Uhr)
🔧 ProgrammierungGitHub Release: langwatch/scenario vjavascript/v1.5.0 (02.09.2026)(02.09.2026 um 04:47 Uhr)
🔧 ProgrammierungGitHub Release: langwatch/scenario vjavascript/v1.6.0 (06.09.2026)(06.09.2026 um 17:45 Uhr)
🔧 Programmierungclawpatrol v0.5.10(13.09.2026 um 02:54 Uhr)
⚠️ Malware / Trojaner / VirenCAPE-parsers v0.1.69(13.09.2026 um 04:14 Uhr)
⚠️ Malware / Trojaner / Virendarknet-mcp-server(13.09.2026 um 04:55 Uhr)
🐧 Linux Tippsazurelinux v3.0.20260909-3.0(13.09.2026 um 09:51 Uhr)
🕵️ Sicherheitslückenatomicvulns(13.09.2026 um 10:36 Uhr)
🔧 ProgrammierungGitHub Release: dependabot/dependabot-core v0.395.0 (07.09.2026)(07.09.2026 um 15:21 Uhr)
🔧 ProgrammierungGitHub Release: dependabot/dependabot-core v0.396.0 (14.09.2026)(14.09.2026 um 19:05 Uhr)
🔧 ProgrammierungGitHub Release: langwatch/scenario vpython/v1.4.0 (02.09.2026)(02.09.2026 um 04:47 Uhr)
🔧 ProgrammierungGitHub Release: langwatch/scenario vjavascript/v1.5.0 (02.09.2026)(02.09.2026 um 04:47 Uhr)
🔧 ProgrammierungGitHub Release: langwatch/scenario vjavascript/v1.6.0 (06.09.2026)(06.09.2026 um 17:45 Uhr)
🔧 Programmierungclawpatrol v0.5.10(13.09.2026 um 02:54 Uhr)
⚠️ Malware / Trojaner / VirenCAPE-parsers v0.1.69(13.09.2026 um 04:14 Uhr)
⚠️ Malware / Trojaner / Virendarknet-mcp-server(13.09.2026 um 04:55 Uhr)
🐧 Linux Tippsazurelinux v3.0.20260909-3.0(13.09.2026 um 09:51 Uhr)
🕵️ Sicherheitslückenatomicvulns(13.09.2026 um 10:36 Uhr)

🔧 Programmierung 🕛 vor 3 Monaten 33 Min Lesezeit
0

Build a Production RAG System on AWS Bedrock from Scratch

↗ Quelle (dev.to)
🗣️ Stimme:
📑 Inhaltsübersicht

A complete hands-on guide to building Retrieval Augmented Generation on AWS Bedrock with pgvector, Guardrails, Prompt Management, Knowledge Bases, Evaluations and API Gateway





Building and running this system teaches you more than reading about it. The specific errors you hit (like the bedrock-agent-runtime missing VPC endpoint causing a 5-minute Lambda timeout, or RetrieveAndGenerate rejecting cross-region inference profile ARNs) are exactly the kind of edge cases the exam tests.






Architecture



Architecture Diagram showing all AWS services and two data flows: Document Ingest and Query]



The diagram above shows two distinct flows:



Document Ingest Flow (blue):




  1. A file is uploaded to S3 under the docs/ prefix

  2. S3 event notification triggers the Ingest Lambda

  3. Lambda reads the file, chunks it (800 tokens, 100 token overlap)

  4. Each chunk is embedded via Bedrock Titan Embeddings v2 (1024 dimensions)

  5. Embeddings are upserted into Aurora Serverless v2 with pgvector (HNSW index)



Query Flow (green):




  1. Client sends POST /query with a JWT Bearer token

  2. API Gateway validates the JWT against Cognito

  3. Lambda embeds the question with Titan

  4. Lambda runs vector similarity search in Aurora pgvector (top-5 results)

  5. Lambda fetches the versioned prompt template from Bedrock Prompt Management

  6. Lambda calls Claude Haiku 4.5 with the context, applying Guardrails

  7. Response is returned with sources, similarity scores, and the prompt ARN used

  8. Conversation turn is saved to DynamoDB for session history



Knowledge Base Flow (orange): An alternative path via POST /query-kb routes to Bedrock's managed RetrieveAndGenerate API, completely skipping the custom pgvector retrieval.






Why a VPC With No NAT Gateway?



All compute runs in private subnets with zero internet access. Every AWS service call goes through VPC endpoints (PrivateLink), which means:




  • No data ever traverses the public internet

  • No NAT Gateway cost (~£30/month saving)

  • Traffic to S3 and DynamoDB uses free gateway endpoints

  • Five interface endpoints handle Bedrock, Secrets Manager, and CloudWatch Logs



This is a realistic enterprise configuration and a common exam topic.






Prerequisites



Before starting you need:




  • AWS account with admin IAM user credentials configured via aws configure

  • Python 3.12 installed locally

  • Git



Add Marketplace permissions to your IAM user. The newer Claude models (Haiku 4.5, the EU cross-region inference profiles) require your IAM user to have these permissions to invoke them the first time:




  1. IAM console → Users → your user → Add permissionsCreate inline policy

  2. Choose JSON editor and paste:



CODE
{
"Version": "2012-10-17",
"Statement": [{
"Effect": "Allow",
"Action": [
"aws-marketplace:ViewSubscriptions",
"aws-marketplace:Subscribe",
"aws-marketplace:Unsubscribe"
],
"Resource": "*"
}]
}


3. Name the policy bedrock-marketplace and save.




AIP-C01 note: Cross-region inference profiles (the eu. prefix on a model ID) route requests across AWS regions for higher availability. The first invocation per AWS account requires Marketplace subscription by the invoking principal. Direct foundation model IDs (e.g. anthropic.claude-3-7-sonnet-20250219-v1:0) do not require this.







Phase 1: Supporting Resources






1.1 Create the S3 Docs Bucket




  1. S3 console → Create bucket

  2. Bucket name: rag-bedrock-docs-xxx

  3. Region: eu-west-2

  4. Block all public access: enabled (leave default)

  5. Versioning: Enable

  6. Default encryption: SSE-S3

  7. Create bucket






  1. DynamoDB console → Create table

  2. Table name: rag-bedrock-sessions

  3. Partition key: session_id (String)

  4. Sort key: timestamp (Number)

  5. Table settings: Customize settings

  6. Capacity mode: On-demand

  7. Encryption: Owned by Amazon DynamoDB

  8. Create table



After creation, enable TTL:




  1. Open the table → Actions → Turn on TTL

  2. Type expires_at in the TTL attribute name, leave the preview simulation as-is, then click Turn on TTL.




The expires_at field is what the Lambda writes as a Unix epoch timestamp (current time + 30 days in seconds). DynamoDB will automatically delete session items older than 30 days.



Why DynamoDB for sessions: Lambda is stateless. Each invocation needs to load the last N conversation turns to maintain context. DynamoDB with TTL gives you automatic expiry (30 days), millisecond reads, and no server to manage. The sort key on timestamp lets you query the N most recent turns efficiently.







Phase 2: VPC and Networking



This is the most console-intensive phase. Take your time getting the VPC right means no debugging later.






2.1 Create the VPC




  1. VPC console → Your VPCs → Create VPC

  2. Resources to create: VPC and more

  3. Name tag: rag-bedrock

  4. IPv4 CIDR: 10.42.0.0/16

  5. Number of Availability Zones: 2 (eu-west-2a, eu-west-2b)

  6. Number of public subnets: 0 (no public subnets needed)

  7. Number of private subnets: 2

  8. NAT gateways: None (we use VPC endpoints instead — saves ~£30/month)

  9. VPC endpoints: None (we create them manually below for full control)

  10. Leave everything else to default.

  11. Create VPC








2.3 Create VPC Endpoints



You need five interface endpoints and two gateway endpoints. Create each via VPC console → Endpoints → Create endpoint.



Gateway endpoints (free create these first):



Endpoint 1 S3:




  • Service category: AWS services

  • Service name: search for s3 and select com.amazonaws.eu-west-2.s3 (Gateway type)

  • VPC: your VPC

  • Route tables: select the private route tables for both subnets



Endpoint 2 DynamoDB:




  • Service name: com.amazonaws.eu-west-2.dynamodb (Gateway type)

  • Same VPC and route tables



Interface endpoints (billable at ~£0.008/hr/AZ each):



For each of the five below, use:




  • Service category: AWS services

  • VPC: your VPC

  • Subnets: both private subnets

  • Security group: rag-bedrock-endpoints-sg

  • Private DNS enabled: Yes






AIP-C01 note: VPC endpoint to Bedrock service mapping: A common exam scenario is “a Lambda in a private subnet cannot reach Bedrock Knowledge Bases.” The answer is a missing bedrock-agent-runtime endpoint. Know which endpoint enables which API:



bedrock-runtime → InvokeModel



bedrock-agent → GetPrompt (Prompt Management)



bedrock-agent-runtime → RetrieveAndGenerate and Retrieve (Knowledge Bases)







Phase 3: Aurora Serverless v2 with pgvector






3.1 Create the Aurora Cluster




  1. RDS console → Create database → Full Configuration

  2. Engine: Aurora (PostgreSQL Compatible)

  3. Templates: Dev/Test

  4. Cluster scalability type: Serverless v2

  5. Capacity range: Min 0 ACU, Max 2 ACU (scale-to-zero when idle)

  6. Engine version: Aurora PostgreSQL 17.7 (or latest )

  7. DB cluster identifier: rag-bedrock-cluster

  8. Credentials Settings: Master username ragdb

  9. Credentials: Managed in AWS Secrets Manager (check this box — it auto-creates and rotates the password)

  10. Connectivity:




  • VPC: your VPC

  • DB subnet group: create a new DB subnet group using both private subnets

  • Public access: No

  • VPC security group: rag-bedrock-aurora-sg

  • Enable the RDS Data API checkbox (required for the schema bootstrap)



11. Additional configuration:




  • Database port: 5432

  • Initial database name: ragdb



12. Encryption: AWS managed key (do NOT use a customer-managed key — if you destroy the cluster, a CMK goes into PendingDeletion for 7–30 days and makes automated snapshots unrestorable)



13. Leave everything else to default.



14. Create database



.Wait for the cluster status to show Available (3–5 minutes).






3.2 Bootstrap the Database Schema



Aurora is now running but has no tables. Use the RDS Query Editor (built into the console no VPN or bastion needed) to run the setup SQL.




  1. RDS console → Query Editor

  2. Cluster: rag-bedrock-cluster

  3. Database: ragdb

  4. Authentication: Connect with a Secrets Manager ARN → paste your Secret ARN

  5. Run the following SQL statements one at a time:



CODE
-- Enable pgvector extension
CREATE EXTENSION IF NOT EXISTS vector;

-- Source files tracking table
CREATE TABLE IF NOT EXISTS source_files (
id bigserial PRIMARY KEY,
s3_key text NOT NULL UNIQUE,
ingested_at timestamptz DEFAULT now()
);

-- Document chunks with 1024-dimension vectors (Titan v2)
CREATE TABLE IF NOT EXISTS documents (
id bigserial PRIMARY KEY,
source text NOT NULL,
chunk_index integer NOT NULL,
content text NOT NULL,
embedding vector(1024),
metadata jsonb,
created_at timestamptz DEFAULT now(),
UNIQUE(source, chunk_index)
);

-- HNSW index for fast approximate nearest-neighbour search
CREATE INDEX IF NOT EXISTS documents_embedding_idx
ON documents
USING hnsw (embedding vector_cosine_ops)
WITH (m = 16, ef_construction = 64);





  1. IAM console → Roles → Create role

  2. Trusted entity: AWS service → Lambda

  3. Attach these managed policies:




  • AWSLambdaBasicExecutionRole

  • AWSLambdaVPCAccessExecutionRole



4. Name the role: rag-bedrock-lambda-role



5. Create role



Now add a custom inline policy. Open the role → Add permissions → Create inline policy →




Note: Replace YOURACCOUNTID with your 12-digit AWS account ID. Name the policy rag-bedrock-lambda-policy.




CODE
{
"Version": "2012-10-17",
"Statement": [
{
"Sid": "BedrockInvoke",
"Effect": "Allow",
"Action": [
"bedrock:InvokeModel",
"bedrock:InvokeModelWithResponseStream",
"bedrock:ApplyGuardrail"
],
"Resource": "*"
},
{
"Sid": "BedrockKB",
"Effect": "Allow",
"Action": [
"bedrock:Retrieve",
"bedrock:RetrieveAndGenerate",
"bedrock:GetPrompt"
],
"Resource": "*"
},
{
"Sid": "SecretsRead",
"Effect": "Allow",
"Action": [
"secretsmanager:GetSecretValue",
"secretsmanager:DescribeSecret"
],
"Resource": "arn:aws:secretsmanager:eu-west-2:YOURACCOUNTID:secret:rds!*"
},
{
"Sid": "S3Docs",
"Effect": "Allow",
"Action": ["s3:GetObject", "s3:PutObject", "s3:ListBucket"],
"Resource": [
"arn:aws:s3:::rag-bedrock-docs-YOURACCOUNTID",
"arn:aws:s3:::rag-bedrock-docs-YOURACCOUNTID/*"
]
},
{
"Sid": "DynamoSessions",
"Effect": "Allow",
"Action": [
"dynamodb:GetItem",
"dynamodb:PutItem",
"dynamodb:Query",
"dynamodb:UpdateItem"
],
"Resource": "arn:aws:dynamodb:eu-west-2:YOURACCOUNTID:table/rag-bedrock-sessions"
},
{
"Sid": "KmsViaSvc",
"Effect": "Allow",
"Action": ["kms:Decrypt", "kms:DescribeKey"],
"Resource": "*",
"Condition": {
"StringEquals": {
"kms:ViaService": [
"secretsmanager.eu-west-2.amazonaws.com",
"rds.eu-west-2.amazonaws.com"
]
}
}
}
]
}





4.2 Package the Lambda Code



You need to create two deployment zip files from the GitHub repo. Run these commands from the repo root:



CODE
# Clone the repo if you haven't already
git clone https://github.com/joysontech/rag-bedrock.git
cd rag-bedrock

# ── Ingest Lambda ──────────────────────────────────────
pip3 install -r src/ingest/requirements.txt \
-t ~/Desktop/lambda-packages/ingest-package
cp src/ingest/handler.py ~/Desktop/lambda-packages/ingest-package/
cp -r src/shared ~/Desktop/lambda-packages/ingest-package/
cd ~/Desktop/lambda-packages/ingest-package && zip -r ~/Desktop/lambda-packages/ingest.zip . && cd ~/rag-bedrock

# ── Query Lambda ───────────────────────────────────────
pip3 install -r src/query/requirements.txt \
-t ~/Desktop/lambda-packages/query-package
cp src/query/handler.py ~/Desktop/lambda-packages/query-package/
cp -r src/shared ~/Desktop/lambda-packages/query-package/
cd ~/Desktop/lambda-packages/query-package && zip -r ~/Desktop/lambda-packages/query.zip . && cd ~/rag-bedrock

# ── Verify both zips contain psycopg ──────────────────
echo "=== ingest ===" && unzip -l ~/Desktop/lambda-packages/ingest.zip | grep pg8000
echo "=== query ===" && unzip -l ~/Desktop/lambda-packages/query.zip | grep pg8000

echo "Done. Files on Desktop:"
ls -lh ~/Desktop/lambda-packages/*.zip



The --platform manylinux2014_x86_64 --only-binary=:all: flags ensure psycopg3 installs the correct Linux binary even when packaging from a Mac or Windows machine.







4.3 Create the Ingest Lambda








4.4 Create the Query Lambda



Repeat the same steps as the Ingest Lambda with these differences:




  • Function name: rag-bedrock-query

  • Upload /tmp/query.zip

  • Same environment variables, plus these additional ones (add them now, update values after creating Guardrails and Prompt Management):








4.6 Upload the exam guide document from the repo



S3 console → your bucket → Upload




  1. Open the S3 console → click rag-bedrock-docs-YOURACCOUNTID

  2. You’ll see the bucket is empty (or has placeholder files). Click Create folder

  3. Folder name: docs → Create folder

  4. Click into the docs/ folder

  5. Click Upload

  6. Click Add files → navigate to your local repo → select docs/aip-c01-exam-guide.md

  7. Leave all other settings as default

  8. Click Upload



The file will be at s3://rag-bedrock-docs-YOURACCOUNTID/docs/aip-c01-exam-guide.md which matches the docs/ prefix filter on the S3 event notification.



Verify it triggered the Lambda:



CODE
aws logs tail /aws/lambda/rag-bedrock-ingest --since 3m --follow




Or Verify in RDS Query Editor:



CODE
SELECT source, count(*) AS chunks
FROM documents
GROUP BY source;


Should show docs/aip-c01-exam-guide.md | 8






  1. Cognito console → Create user pool

  2. Application type: Leave as Traditional web application

  3. Name your application:rag-bedrock-app

  4. Sign-in identifiers: Keep Email checked only

  5. Self-registration: Uncheck “Enable self-registration” you will create the test user manually via CLI, no public sign-up needed

  6. Required attributes: Click the dropdown and select email

  7. Return URL: Type






    5.2 Create the API Gateway HTTP API




    1. API Gateway console → Create API → HTTP API → Build

    2. API name: rag-bedrock-api

    3. Integration: Lambda

    4. AWS region: eu-west-2

    5. Lambda function: rag-bedrock-query

    6. Configure routes: keep defaults for now (you will add routes manually)

    7. Stage name: $default

    8. Auto-deploy: On

    9. Create



    Note the Invoke URL shown after creation this is your API endpoint.



    (replace YOURPOOLID)



    7. Audience: your app client ID



    8. Create






    AIP-C01 note API Gateway HTTP API vs REST API: HTTP APIs are the modern choice for Lambda integrations. They support JWT authorizers natively, have lower latency, cost ~70% less, and support payload format 2.0 (which simplifies the Lambda event structure). REST APIs are needed for API keys, request/response transformations, and WAF integration. For RAG backends, HTTP API is the right choice.







    Phase 6: Bedrock Guardrails



    Guardrails sit between the user’s question and the model. They inspect both input and output and can block, mask, or substitute responses.




    1. Bedrock console → Guardrails → Create guardrail

    2. Name: rag-bedrock-guardrail

    3. Content filters (next page):




    • Set all six categories (Hate, Insults, Sexual, Violence, Misconduct) to Medium for both input and output

    • Set Prompt Attack to High — this is your primary prompt injection defence



    4. Denied topics (next page):




    • Add topic: Name = Personal financial advice

    • Definition: Any advice or recommendations about investments, stocks, funds, or personal financial planning

    • Example phrases: "Should I invest in" and "What stocks should I buy"



    5. Add word filters — optional




    • configure as you like



    6. Sensitive information filters (next page):




    • Email address: Mask

    • Phone number: Mask

    • Credit card number: Block

    • UK National Insurance number: Block



    7. Contextual grounding (next page):




    • Enable grounding filter: Yes, threshold: 0.75

    • Enable relevance filter: Yes, threshold: 0.75



    8. Review and create



    After creation, click Create version to publish version 1.



    Note the Guardrail ID (e.g. hy14n4r45o6f). Now update your Query Lambda environment variables:




    • GUARDRAIL_ID = your Guardrail ID

    • GUARDRAIL_VERSION = 1








    Phase 7: Bedrock Prompt Management



    Prompt Management versions your system prompt like code. Instead of hardcoding it in Lambda, it lives in Bedrock with a version history and an audit trail in every API response.




    1. Bedrock console → Prompt management → Create prompt

    2. Name: rag-query-generate

    3. Description: RAG generation prompt forAIP-C01 exam Q&A

    4. Model: Claude Haiku 4.5 (EU Anthropic Claude Haiku 4.5)

    5. Temperature: 0.2, Max tokens: 1024

    6. System instructions (optional field):



    CODE
    You are a helpful assistant answering questions using only the provided context. Never use outside knowledge. Cite sources inline as [source-key].


    7. User message template (use {{double_braces}} for variables):



    CODE
    Context:
    {{context}}

    User question: {{question}}

    Answer using only the context above. If the answer is not in the context, say "I don't have enough information to answer that." Cite sources inline as [source-key].Important: if an “Assistant message” field appears, delete it using the trash icon. An empty assistant message block causes a ContentBlock is blank validation error.


    8. Important: if an “Assistant message” field appears, delete it using the trash icon. An empty assistant message block causes a ContentBlock is blank validation error.



    9. In the Test variables section:



    context:



    CODE
    [Source: docs/aip-c01-exam-guide.md] The AIP-C01 exam is divided into five domains. Domain 1: Foundation Model Integration and Data Management covers 31% of the exam. Domain 2: GenAI Application Implementation and Integration covers 26%. Domain 3: AI Safety, Security and Governance covers 20%. Domain 4: Operational Excellence and Efficiency covers 12%. Domain 5: Testing, Validation and Troubleshooting covers 11%.


    question:



    CODE
    What percentage of the AIP-C01 exam does Domain 1 cover?


    10. Click Run to verify the prompt returns a grounded answer





    After creation:




    1. Click Sync on the data source to trigger initial ingestion

    2. Note the Knowledge Base ID (e.g. TMBSW0OWMK)



    Update your Query Lambda environment variable:




    • KNOWLEDGE_BASE_ID = your Knowledge Base ID







    Both systems return factually correct, grounded answers. The difference is observability. DIY tells you exactly what was retrieved, how similar it was, which prompt version generated the answer, and what the user asked before. Knowledge Base gives you the answer with no visibility into the why.




    AIP-C01 note — Retrieve vs RetrieveAndGenerate: The Knowledge Base has two API operations:





    • Retrieve: returns relevant chunks with scores. Use this when you want to control generation yourself (custom model, Guardrails, Prompt Management).

    • RetrieveAndGenerate: end-to-end RAG in one call. Use this when you want simplicity over control.




    This system uses RetrieveAndGenerate for the /query-kb route. If you wanted to apply Guardrails to the KB path too, you would use Retrieve + InvokeModel instead.







    Phase 9: Bedrock Evaluations



    Evaluations measure whether your system is actually producing good answers using LLM-as-judge: a stronger model (Claude Sonnet) evaluates outputs from your generation model (Claude Haiku).






    9.1 Upload the Evaluation Dataset



    Create a JSONL file with question/reference-answer pairs:



    CODE
    {"prompt": "Which VPC endpoint enables the RetrieveAndGenerate API for Knowledge Bases?", "referenceResponse": "The bedrock-agent-runtime VPC endpoint enables the RetrieveAndGenerate and Retrieve APIs for Bedrock Knowledge Bases.", "category": "question_answering"}
    {"prompt": "What percentage of the AIP-C01 exam does Domain 1 cover?", "referenceResponse": "Domain 1 (Foundation Model Integration and Data Management) covers 31% of the AIP-C01 exam.", "category": "question_answering"}
    {"prompt": "When should I use RAG instead of fine-tuning?", "referenceResponse": "Use RAG for frequently updated knowledge, large document corpora, and when auditability matters. Use fine-tuning for style consistency, domain vocabulary, and classification tasks with static training data.", "category": "question_answering"}
    {"prompt": "What is the default chunking strategy in Bedrock Knowledge Bases?", "referenceResponse": "The default chunking strategy in Bedrock Knowledge Bases is 300 tokens with 20% overlap. The chunking strategy cannot be changed after creating a data source.", "category": "question_answering"}
    {"prompt": "Which AWS service should I use as a vector store for large-scale RAG with hybrid search support?", "referenceResponse": "Amazon OpenSearch Serverless supports both vector search and keyword search (hybrid search), making it the recommended choice for large-scale RAG deployments requiring hybrid search.", "category": "question_answering"}
    {"prompt": "What is the passing score for the AIP-C01 exam?", "referenceResponse": "The passing score for the AIP-C01 exam is 720 out of 1000.", "category": "question_answering"}
    {"prompt": "What is the difference between the Retrieve and RetrieveAndGenerate Knowledge Base APIs?", "referenceResponse": "Retrieve returns relevant document chunks with scores for you to control generation yourself. RetrieveAndGenerate handles both retrieval and generation in one call for simpler Q&A without custom generation control.", "category": "question_answering"}
    {"prompt": "What happens if a Lambda in a private VPC is missing the bedrock-agent-runtime endpoint?", "referenceResponse": "Without the bedrock-agent-runtime VPC endpoint, the Lambda call to Bedrock Knowledge Bases silently hangs for the full timeout duration rather than returning an immediate error.", "category": "question_answering"}


    Save as eval-dataset.jsonl and upload to S3: via console



    Or CLI



    CODE
    aws s3 cp eval-dataset.jsonl \
    s3://rag-bedrock-docs-YOURACCOUNTID/evals/eval-dataset.jsonl





    9.2 Create the Evaluation Job




    1. Bedrock console → Evaluations → Create → Automatic: LLM as a judge

    2. Job name: rag-bedrock-eval-v1

    3. Evaluator (judge) model: Claude Sonnet 4.6 (stronger than the evaluated model this is the LLM-as-judge pattern)

    4. Inference source: Bedrock models

    5. Generator model (model being evaluated): Claude Haiku 4.5 EU inference profile

    6. Metrics — tick these four (deselect the defaults):




    • Correctness: is the answer factually right versus the reference?

    • Faithfulness: is the answer grounded in context, not hallucinated?

    • Completeness: does it fully answer the question?

    • Relevance: is it on-topic?



    7. Dataset S3 URI: s3://rag-bedrock-docs-YOURACCOUNTID/evals/eval-dataset.jsonl



    8. Output S3 URI: s3://rag-bedrock-docs-YOURACCOUNTID/evals/results/



    9. IAM role: Create and use a new service role



    10. Create



    The job takes 5–10 minutes. Check results in the Bedrock Evaluations console once it completes.






    AIP-C01 note — LLM-as-judge: The judge should be a stronger model than the evaluated model. Here Sonnet (stronger) judges Haiku (weaker). The referenceResponse is the gold-standard answer. The judge scores the model's actual response against it on each metric, returning a score from 0 to 1. Scores near 1 on faithfulness mean the model is not hallucinating — everything it says is traceable to the retrieved context.







    Errors I Hit and How to Fix Them



    These are the real errors encountered when building this system. Every one of them will appear in some form when you follow this guide.






    Error 1: Runtime.ImportModuleError: No module named 'lambda_function'



    When: Lambda invoked for the first time after uploading the zip.



    Why: The Lambda console defaults the handler to lambda_function.lambda_handler. The code in this repo uses handler.handler (file: handler.py, function: handler).



    Fix: Lambda console → your function → Code tab → scroll down to Runtime settings → Edit → change Handler to handler.handler → Save.






    Error 2: Runtime.ImportModuleError: no pq wrapper available (psycopg3)



    When: Lambda starts after uploading a zip built with psycopg[binary].



    Why: The psycopg3 binary wheel (psycopg-binary) is not available for the manylinux_2_28_x86_64 platform used by Lambda Python 3.12. Packaging from a Mac downloads an incompatible binary.



    Fix: This repo now uses pg8000 — a pure Python PostgreSQL driver with no binary dependencies. No platform flags are needed when packaging:



    CODE
    pip3 install -r src/ingest/requirements.txt \
    -t ~/Desktop/lambda-packages/ingest-package





    Error 3: Runtime.ImportModuleError: No module named 'psycopg2._psycopg'



    When: Lambda starts after uploading a zip built with psycopg2-binary on Mac.



    Why: The psycopg2-binary wheel downloaded with --platform manylinux2014_x86_64 contains a .so file compiled for a different Python version or glibc version than Lambda's runtime. Packaging with --python-version 3.12 --implementation cp helps but can still fail depending on the version.



    Fix: Use pg8000 (pure Python, no compilation, no platform flags). The repo requirements files use pg8000==1.31.2.






    Error 4: AccessDeniedException: not authorized to perform s3:GetObject



    When: Ingest Lambda triggers on S3 upload but fails at step 1.



    Why: The IAM inline policy on the Lambda role was created with rag-bedrock-docs-YOURACCOUNTID as a placeholder. The actual bucket name was different (e.g. rag-bedrock-docs-demo123).



    Fix: IAM console → your Lambda role → inline policy → edit → replace the placeholder bucket name with your actual bucket name in both the bucket ARN and the /* ARN. Save.






    Error 5: KeyError: 'AURORA_SECRET_ARN'



    When: Query Lambda returns Internal Server Error after the first API Gateway call.



    Why: Environment variables were set on the Ingest Lambda but not copied to the Query Lambda. The Query Lambda had no env vars at all.



    Fix: Lambda console → rag-bedrock-query → Configuration → Environment variables → Edit → add all required env vars. See Phase 4.4 for the complete list.






    Error 6: Route tables not appearing when creating S3 gateway endpoint



    When: Creating the S3 gateway endpoint in VPC → Endpoints — the route table dropdown is empty.



    Why: The VPC console wizard does not always create explicit route table associations for the private subnets. The subnets use the main route table implicitly.



    Fix: VPC console → Route tables → Create route table → name it rag-bedrock-private-rt → associate both private subnets → then create the gateway endpoints and select this route table.






    Error 7: NotAuthorizedException: Client configured with secret but SECRET_HASH was not received



    When: Running aws cognito-idp initiate-auth with USER_PASSWORD_AUTH.



    Why: The Cognito app client was created as a confidential client (with a secret). The initiate-auth CLI command does not support computing the SECRET_HASH — that requires additional code.



    Fix: Cognito console → your user pool → App clients → Create app client → choose Public client → toggle Generate client secret OFF → use the new client ID for all CLI commands.






    Error 8: InvalidParameterException: Attributes did not conform to the schema: emails: The attribute emails is required



    When: Running aws cognito-idp sign-up without --user-attributes.



    Why: When email is configured as a required sign-in identifier, Cognito requires the email attribute to be passed explicitly even when it is also used as the username.



    Fix: Add --user-attributes Name=email,Value=



    Total: ~£32/month — almost entirely the VPC endpoints. In a real AWS account you would share those endpoints across multiple services, so the marginal cost for this project is much lower.



    To save costs when not actively testing: scale Aurora to 0 manually (it happens automatically after 5 minutes of idle) and optionally delete the VPC interface endpoints (keeping just the gateway endpoints). Recreating interface endpoints takes 2–3 minutes via the console.






    AIP-C01 Quick Reference






    Bedrock model IDs in eu-west-2



    CODE
    # Direct invocation (no Marketplace required):
    anthropic.claude-3-7-sonnet-20250219-v1:0
    anthropic.claude-sonnet-4-6

    # Cross-region inference profile (requires Marketplace subscription first):
    eu.anthropic.claude-haiku-4-5-20251001-v1:0

    # RetrieveAndGenerate model ARN (NO eu. prefix — inference profiles rejected):
    arn:aws:bedrock:eu-west-2::foundation-model/anthropic.claude-3-7-sonnet-20250219-v1:0


    VPC endpoint → Bedrock API mapping





    For RAG with Titan Embeddings v2, always use <=>.






    Guardrail stop reasons




    • end_turn: normal completion, no intervention

    • guardrail_intervened: hard block, check amazon-bedrock-guardrailAction in response body

    • Substituted message: HTTP 200 but content replaced — check the actual response text






    Chunking strategies (exam topic)





  8. AIP-C01 Udemy course:

  9. pgvector: github.com/pgvector/pgvector



  10. Drop a comment with your eval scores curious how Haiku 4.5 performs on correctness vs faithfulness on different document types.

    Vollständiger Original-Bericht
    Ausführliche Details, Code-Beispiele & Hersteller-Stellungnahme auf dev.to.
    ↗ Original-Artikel auf dev.to lesen
Wie bewertest du diesen Beitrag?
1 Klick Feedback
Teilen mit Netzwerk & Team:

Community-Analysen & Experten-Meinungen 0

Verfasse deine eigene Analyse, teile Workarounds oder diskutiere diesen Vorfall im Blog.
Noch keine Community-Analyse verfasst. Markiere einen Textabschnitt oder klicke oben auf Eigene Analyse verfassen“!
Community Pulse: Relevanz-Einschätzung
1 Klick Experten-Votum
🔴 Akute Relevanz 0%
🟡 In Evaluierung 0%
🟢 Keine Auswirkung 0%
Spannende Innovation 0%
Verwandte Story-Cluster & Quellen (Vektor-KI)
Port 8095 Engine
10 Quellen
GitHub Release: dependabot/dependabot-core v0.393.0 (24.08.2026)
1 Quelle
clawpatrol v0.5.10
1 Quelle
CAPE-parsers v0.1.69
Ähnliche Beiträge
🔍 Verwandte News

Auch interessante Nachrichten Build a Production RAG System on AWS Bedrock from Scratch

Thematisch verwandte Begriffe: Build, Production, System, Bedrock · 6 Treffer

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...