💬 I don’t really know Python.
My background is Pascal, VB, and Prolog — structured and logical, but far from modern languages.
This system was built under pressure, pushed forward by Cursor and Copilot.
This post is about how I designed a self-correcting image prompt generator using a multi-stage LLM flow.
🎯 Why One Prompt Isn’t Enough
Most LLM prompt systems work like this:
- You give it a caption
- It generates tags or scene descriptions
- You hope the structure makes sense
But most of the time, it doesn’t.
In a system like RΞNE — where consistency matters (we feed prompts to Stable Diffusion) — I didn’t want to manually audit outputs. So I built /genimgprompt: a self-checking, retry-capable multi-stage pipeline.
🧠 Step 1: Caption → Scene (Creative Persona)
We begin with caption + emotion + character + style_hint. The first stage uses a creative persona (RΞNE) to generate a vivid scene.
Example input:
{
"caption": "A cyberpunk girl walks through a neon alley",
"emotion": { "primary": "MYSTERIOUS", "intensity": 0.7 },
"character": { "desc": "A mysterious girl with cybernetic eyes" },
"style_hint": "cyberpunk, neon, atmospheric"
}
Scene result:
“She walks alone through the flickering neon-lit alley, her cybernetic eyes faintly glowing, casting light onto the damp ground. The atmosphere is saturated with mystery, her silhouette melting into the blurred textures of the cyberpunk city.”
🧩 Step 2: Scene → Structured Tags (Tagify)
Next, we ask the LLM to extract 6 structured fields:
characterpose_actionoutfitemotionbackgroundcamera
Then we parse:
If the response fails parsing, we trigger retry logic:
- Retry with
AI_Assistantpersona (more structured) - Still invalid? Use a default fallback
🔁 Retry and Fallback Mechanism
Here's the logic:
try:
response = call_llm(persona="RΞNE")
tags = parse_tagify(response)
except ParseError:
response = call_llm(persona="AI_Assistant")
tags = parse_tagify(response)
if not valid(tags):
tags = default_prompt_tags()
Why this matters:
- LLMs are unreliable no matter how strict your prompt is
- Any slight formatting deviation breaks downstream logic
- Debug info is logged at every stage
🧰 What I’m Open Sourcing
I’ll open source a simplified FastAPI module with:
- Input:
caption,emotion,character,style_hint
- Output:
prompt_tags,positive_prompt,negative_prompt
- Includes full debug logs of retries, LLM calls, and fallback triggers
- No memory system, no Stable Diffusion dependency, no image output
This module is fully standalone and can be integrated into any AI generation pipeline.
💭 Final Thoughts
This module isn’t about perfect code — it’s about a system that fails gracefully, recovers, and tells you what happened.
Side note: using AI/LLMs for coding is one of the best things that’s happened to me.
📡 Follow @n40-rene.bsky.social
Next post: full code and open-source repo release.