Zum Hauptinhalt springen
tsecurity.de LIVE
Echtzeit-Radar & Feeds
Alle RSS Feeds
👥 Community & Social
IT NachrichtenPeloton has made a foldable (treadmill)(22.09.2026 um 13:00 Uhr)
IT NachrichtenPeloton has made a foldable (treadmill)(22.09.2026 um 13:00 Uhr)
Intelligence View
⚡ tsecurity.de Intelligence

Building a Ride-Share Zone-Balancing Agent with LangGraph — Part 3: Giving the Agent Memory

This is Part 3 of a 5-part series. Part 1 built a rule-based agent for one zone. Part 2 added an LLM for two narrow jobs: reading an ops note, and explaining the decision. Neither part changed one basic fact: the agent looks at a zone…

0
↗ Quelle (dev.to)
Reagiere als Erste:r — dein Feedback zählt!

This is Part 3 of a 5-part series. Part 1 built a rule-based agent for one zone. Part 2 added an LLM for two narrow jobs: reading an ops note, and explaining the decision. Neither part changed one basic fact: the agent looks at a zone exactly once, then stops.



That's a real problem. A severe deficit often needs more than one 15-minute window to fix. But calling the agent again is first contact all over again:




  • No memory of what was already tried

  • No sense of whether this is cycle one or cycle six

  • The zone's own numbers reset, instead of carrying forward



Part 3 fixes that with MemorySaver and a thread_id. One thread becomes one zone's ongoing story. One .invoke() call becomes one 15-minute cycle.



The decision logic itself doesn't change at all. detect_imbalance, classify_severity, resolved_imbalance, and choose_best_policy are the exact same deterministic functions from Part 1, reused directly.



What memory actually buys here:




  • The zone's real state carries forward correctly, instead of resetting. That includes drivers, riders, and a little organic replenishment each cycle.

  • A note can schedule something for future cycles — "traffic bad for the next hour," "the event starts at 6pm" — without being repeated every time. The LLM extracts the fact once. Code applies it automatically from then on.






The Graph



Two new nodes get added at the front. Both are plain Python — no LLM:




start_cycle ──▶ apply_scheduled_conditions ──▶ detect_imbalance ──▶ classify_severity ──▶ set_candidates

┌────────────┴────────────┐
▼ (balanced) ▼ (deficit/surplus)
trivial_do_nothing reconcile_inputs ← LLM #1
│ ▼
│ resolved_imbalance
│ ▼
│ choose_best_policy
│ ▼
│ generate_explanation ← LLM #2
│ │
└──────────────┬───────────┘

simulate_and_report








  • start_cycle clears the previous cycle's raw conversation. This keeps the message log from growing forever across separate .invoke() calls. (History still persists — it's kept in its own history field, not in the message log.) It also advances the clock, so demand and supply use the right hour's baseline.


  • apply_scheduled_conditions checks whether anything a past note scheduled should expire or trigger this cycle. No LLM call — just comparing the current cycle number against a stored expiry or trigger.



reconcile_inputs also gets extended, not replaced. It's the same single-shot call from Part 2. Now it also picks up a stated duration ("for the next hour") or a future trigger time ("starts at 6pm"), when the note mentions one.



The LLM extracts the raw fact once. apply_scheduled_conditions applies it automatically on every cycle after that.






The State



AgentState grows four fields past Part 2's:





  • initial_hour and cycle_number anchor the clock


  • history accumulates one record per cycle


  • scheduled_conditions holds whatever a past note scheduled for later



None of this is MemorySaver itself — it's just ordinary state. MemorySaver is what makes this state survive between separate calls, instead of resetting every time:




class AgentState(TypedDict):
zone: dict
ops_note: str
initial_hour: int # set once on cycle 1, anchors the clock
cycle_number: int
history: Annotated[list[dict], operator.add]
scheduled_conditions: dict # e.g. {"traffic": {"expires_at_cycle": 5}}
imbalance_ratio: float
imbalance_type: str
severity: str
candidate_policies: list
policy_evaluations: dict
policy_resolutions: dict
recommended_policy: str
explanation: str
messages: Annotated[list[BaseMessage], add_messages]
outcome: dict
outcome_delta: dict
report: str









The Extended Tool



report_context_and_schedule is Part 2's report_context tool, with two fields added. That lets the LLM report timing, not just a current condition — a stated duration, or a future trigger hour.



Code still owns all the arithmetic that follows: minutes to cycles, hour to cycles-until. The LLM only extracts the raw fact:




@tool
def report_context_and_schedule(
rain_flag: bool,
event_flag: bool,
traffic_level: Literal["none", "light", "moderate", "heavy"],
traffic_duration_minutes: int = 0,
event_starts_at_hour: int = -1,
) -> str:
"""
Call this exactly once with your interpretation of the ops note for THIS
cycle. If the note doesn
't affect a value, report the zone's current value
unchanged.

traffic_duration_minutes: ONLY if the note states how long this traffic
condition will last (e.g.
"for the next hour" -> 60). 0 if no duration is
stated -- traffic_level then applies to this cycle only.

event_starts_at_hour: ONLY if the note describes an event scheduled for a
specific clock time that hasn
't started yet (e.g. "starts at 6pm" -> 18).
-1 if no future start time is stated.
"""
return (
f"rain_flag={rain_flag}, event_flag={event_flag}, traffic_level={traffic_level}, "
f"traffic_duration_minutes={traffic_duration_minutes}, event_starts_at_hour={event_starts_at_hour}"
)









Wiring In the Checkpointer



Every node here is a plain function, same as before. The one genuinely new piece is MemorySaver, passed in at compile() time.



That's what turns zone/history/scheduled_conditions from ordinary state into something that survives across separate .invoke() calls on the same thread_id:




llm = ChatOllama(model="qwen2.5:14b", temperature=0)
llm_with_reconcile_tool = llm.bind_tools([report_context_and_schedule])

def _reconcile(state):
return reconcile_inputs(state, llm_with_reconcile_tool)

def _explain(state):
return generate_explanation(state, llm)

g = StateGraph(AgentState)

g.add_node("start_cycle", start_cycle)
g.add_node("apply_scheduled_conditions", apply_scheduled_conditions)
g.add_node("detect_imbalance", detect_imbalance)
g.add_node("classify_severity", classify_severity)
g.add_node("set_candidates", set_candidates)
g.add_node("trivial_do_nothing", trivial_do_nothing)
g.add_node("reconcile_inputs", _reconcile)
g.add_node("resolved_imbalance", resolved_imbalance)
g.add_node("choose_best_policy", choose_best_policy)
g.add_node("generate_explanation", _explain)
g.add_node("simulate_and_report", simulate_and_report)

g.add_edge(START, "start_cycle")
g.add_edge("start_cycle", "apply_scheduled_conditions")
g.add_edge("apply_scheduled_conditions", "detect_imbalance")
g.add_edge("detect_imbalance", "classify_severity")
g.add_edge("classify_severity", "set_candidates")
g.add_conditional_edges("set_candidates", route_llm_or_skip, {
"trivial_do_nothing": "trivial_do_nothing",
"reconcile_inputs": "reconcile_inputs",
})
g.add_edge("reconcile_inputs", "resolved_imbalance")
g.add_edge("resolved_imbalance", "choose_best_policy")
g.add_edge("choose_best_policy", "generate_explanation")
g.add_edge("generate_explanation", "simulate_and_report")
g.add_edge("trivial_do_nothing", "simulate_and_report")
g.add_edge("simulate_and_report", END)

app = g.compile(checkpointer=MemorySaver())






app here is reused for everything in the rest of this article. One graph, many independent zone stories — told apart by thread_id, not by rebuilding anything.






Running Multiple Cycles



MemorySaver plus a thread_id is the whole mechanism. Call .invoke() again on the same thread_id, and the graph picks up exactly where the last call left off.



Only the first call needs a full initial state. Every call after that only needs whatever's new — an ops note, or nothing at all.



Downtown Core, starting in deficit, no ops notes:




[Cycle 1] ratio=1.93 | policy=surge_pricing | wait 6.3min → 4.6min | resolved=NO
Highest profit at $230.78, even though it did not resolve the imbalance.
[Cycle 2] ratio=1.41 | policy=surge_pricing | wait 4.6min → 4.3min | resolved=YES
RESOLVED






Two calls, same thread_id. The second one already knows this is cycle 2 of an ongoing deficit — nothing had to be passed back in by hand.






Scheduling a Condition Across Cycles



An ops note on cycle 1 says traffic will be bad for the next hour — that's 4 cycles. reconcile_inputs extracts traffic_duration_minutes=60. Code converts that to 4 cycles, and schedules the expiry.



No further notes are needed. apply_scheduled_conditions keeps traffic_level="heavy" active through cycle 4, then reverts it automatically at cycle 5 — exactly on schedule:




notes = {1: "There's a bad accident on the main road backing up traffic for the next hour."}
results = run_cycles(app, "downtown-traffic-duration", downtown, ops_notes=notes)









cycle=1 traffic=1.00 scheduled={'traffic': {'expires_at_cycle': 5}}
cycle=2 traffic=1.00 scheduled={'traffic': {'expires_at_cycle': 5}}
cycle=3 traffic=1.00 scheduled={'traffic': {'expires_at_cycle': 5}}
cycle=4 traffic=1.00 scheduled={'traffic': {'expires_at_cycle': 5}}
cycle=5 traffic=0.00 scheduled={}






One note on cycle 1 shapes five cycles of behavior. Nobody has to remind the agent about the traffic on cycles 2 through 4.






A Future Trigger, and Overriding a Schedule



Two more cases worth showing:



A future trigger, not an immediate correction. "The concert starts at 6pm" — it's currently 5pm. That means event_flag should stay False for now, and flip to True only once that hour actually arrives.



reconcile_inputs extracts event_starts_at_hour=18. Code works out how many cycles away that is, and schedules the flip:




cycle=1 hour=17 event_flag=False
cycle=2 hour=17 event_flag=False
cycle=3 hour=17 event_flag=False
cycle=4 hour=17 event_flag=False
cycle=5 hour=18 event_flag=True
cycle=6 hour=18 event_flag=True






A later note overriding an earlier schedule. Say a new note arrives before a scheduled condition would have expired on its own — "actually the accident's been cleared." It replaces the schedule entirely, instead of waiting for the original timer. New information wins:




notes = {
1: "Traffic is bad for the next hour due to an accident.",
3: "Actually the accident has been cleared, traffic is back to normal.",
}









cycle=1 traffic=1.00 scheduled={'traffic': {'expires_at_cycle': 5}}
cycle=2 traffic=1.00 scheduled={'traffic': {'expires_at_cycle': 5}}






By cycle 2 the zone had already resolved on its own. So the schedule never actually got overridden here. But the mechanism — a new note wiping and replacing an existing schedule — is the same one that fires whenever a correction genuinely arrives mid-schedule.






Does the Policy Actually Change?



Every example so far stops at the first resolution. That hides an important question: does the agent keep surging forever, or does it actually back off once a zone balances?



Running 20 cycles straight through, without stopping early, answers that directly:




cycle=1  type=deficit   ratio=1.93  policy=surge_pricing   drivers=16  riders=18  wait= 3.7min
cycle=2 type=balanced ratio=1.12 policy=do_nothing drivers=15 riders=19 wait= 4.4min
cycle=4 type=deficit ratio=1.36 policy=surge_pricing drivers=14 riders=15 wait= 3.8min
cycle=8 type=deficit ratio=1.31 policy=surge_pricing drivers=17 riders=18 wait= 3.5min
cycle=12 type=deficit ratio=1.46 policy=surge_pricing drivers=14 riders=14 wait= 3.3min
cycle=20 type=balanced ratio=1.08 policy=do_nothing drivers=11 riders=12 wait= 3.6min






Over the full 20 cycles, Downtown Core moves between deficit (surge_pricing) and balanced (do_nothing) four separate times. Not once and done — a genuine, ongoing back-and-forth. Ordinary drift periodically pushes the ratio back over the deficit line, before do_nothing settles it again.



Driver and rider counts stay in a stable, bounded range the whole time — drivers 11–17, riders 12–21. Not drifting, not piling up.






What Changed From Part 2

































Part 2 Part 3
Scope One call, one decision A sequence of cycles on the same zone, via thread_id
State Stateless — every call is first contact
MemorySaver carries zone/history/scheduled_conditions forward automatically
Policy selection Deterministic, reused from Part 1 Same, unchanged
Ops note Corrects this decision only Can also schedule a duration or future trigger across cycles


The decision itself never got smarter. choose_best_policy is identical to Part 1.



What changed: the agent can now tell "still working on it, three cycles in" apart from "first time we've looked at this." And a note like "traffic bad for the next hour" doesn't need to be repeated four times.



What's still missing: every decision here still executes unattended, with nobody checking it first. That's where Part 4 picks up.






Code for this series: github.com/ebiarian/zone-balancing-ridesharing-langgraph-agent



Next — Part 4: a human-in-the-loop pause with interrupt(), so a risky or unusual decision doesn't just run without a second opinion.

Zum Aktualisieren ziehen
ZERO-DAY CVE-2026-94493 | A vulnerability was detected in Gigatech PDV5701 1.0.31_240305_112640. T…
Advisory →
TTS Reader • tsecurity.de Voice
tsecurity.de Icon
tsecurity.de App
Offline-Lesen, Eilmeldungen & 0ms Ladezeit

Installiere tsecurity.de direkt auf deinen Home-Bildschirm für das ultimative Vollbild-Magazinerlebnis ohne Browser-Leisten.

Nächster Beitrag
Themen-Radar & Intelligence Matrix
Echtzeit-Taxonomie nach Angriffsvektoren & Plattformen

tsecurity.de Live Threat Radar

🔴 LIVE RADAR
MONITORING
AKTIV
CVE-DATENBANK
LIVE
🔍
Community Radar & Live Chat
Sentinel Bot online • Live-Stream
Dein Cluster: Security Explorer
Match:
lädt…
Verbindung zum Community-Stream wird aufgebaut...
Bearbeitungsmodus — Senden überschreibt deine Nachricht
Community-Puls — was gerade passiert
lädt…
Aktivitäten deiner Analysten
lädt…
Neues Thema oder Eilmeldung einreichen

Reiche interessante Links, Zero-Days oder Debatten ein. Die Community entscheidet per Upvote über die Veröffentlichung.

Heiß diskutierte Einreichungen
🔖 Gespeicherte Artikel
📂 Keine gespeicherten Artikel vorhanden.
Zurück Ziehen Vor
Links: vorheriger Artikel Rechts: nächster Artikel unten: schließen
News NIS-2 Frühwarnung Tier-1 Intel ⏱️ 3 Min vor 10 Min
Artikeldaten werden geladen...

Zurück: vorheriger Vor: nächster
↗ Original-Quelle
Social Reaktionen Deine Reaktion zählt
Einstufung & Relevanz-Poll 0 Stimmen
In sozialen Netzwerken teilen 1-Klick