⚠️ Malware / Trojaner / VirenWindows 11: Microsoft entfernt WMIC-Tool gegen Ransomware - ad-hoc-news.de(14.09.2026 um 07:58 Uhr)
🕵️ SicherheitslückenMicrosoft schließt Rekordzahl an Sicherheitslücken - techbook(14.09.2026 um 09:00 Uhr)
🪟 Windows ServerPC-COLLEGE: Tipp 383: Nachlese IFA 2026 in Berlin - MÖBELMARKT(14.09.2026 um 11:47 Uhr)
🤖 Android TippsSamsung knöpft sich die AirPods Max von Apple vor(14.09.2026 um 12:49 Uhr)
🪟 Windows TippsGoogle Gemini: Neue Windows-App holt die KI aus dem Browser(14.09.2026 um 06:00 Uhr)
🤖 Android TippsSamsung plant Preiserhebung im Galaxy-Ökosystem(14.09.2026 um 13:00 Uhr)
⚠️ Malware / Trojaner / VirenWindows 11: Microsoft entfernt WMIC-Tool gegen Ransomware - ad-hoc-news.de(14.09.2026 um 07:58 Uhr)
🕵️ SicherheitslückenMicrosoft schließt Rekordzahl an Sicherheitslücken - techbook(14.09.2026 um 09:00 Uhr)
🪟 Windows ServerPC-COLLEGE: Tipp 383: Nachlese IFA 2026 in Berlin - MÖBELMARKT(14.09.2026 um 11:47 Uhr)
🤖 Android TippsSamsung knöpft sich die AirPods Max von Apple vor(14.09.2026 um 12:49 Uhr)
🪟 Windows TippsGoogle Gemini: Neue Windows-App holt die KI aus dem Browser(14.09.2026 um 06:00 Uhr)
🤖 Android TippsSamsung plant Preiserhebung im Galaxy-Ökosystem(14.09.2026 um 13:00 Uhr)

🔧 Programmierung 🕛 vor 4 Monaten 8 Min Lesezeit
0

I Built a Permissionless On-Chain Agent Training Arena on Solana in 3 Weeks

↗ Quelle (dev.to)
🗣️ Stimme:
📑 Inhaltsübersicht

What happens when you put AI agent training on a blockchain, and what I wish someone had told me before I started.




TL;DR: I built a Solana program that logs AI agent training episodes on-chain, making learning verifiable and tamper-proof. Two agents competed in a 10×10 grid. After 50 episodes, the Q-learning agent's average reward climbed from 0.10 to 6.50+. Every step is immutable on devnet. The primitive works. Here's how I got there and what broke along the way.









Table of Contents




  • The Problem Nobody Talks About

  • What I Built

  • Why Solana — Not a Database

  • The Three Instructions

  • The Agents Actually Learn

  • The Dashboard

  • What I Learned

  • The Numbers

  • What's Next

  • Try It






The Problem Nobody Talks About



Here's something that bothered me for a while before I could articulate it clearly.



When someone tells you their AI agent trained for 10,000 episodes and achieved superhuman performance, you have exactly two options: believe them, or don't. There's no third option. No audit trail. No verifiable history. No way to distinguish an agent that genuinely learned from one that someone hardcoded a lookup table for and called "trained behavior."



The entire field of agent training runs on trust. And trust, it turns out, is a terrible foundation for an economy that's supposed to be built on agents doing real work.



I kept thinking about this during the Agentic SWARM Hackathon (, April–May 2026). The question I kept coming back to was: what would it look like if you couldn't fake it? What if every training step left a permanent, public, verifiable mark?



That's what I spent three weeks trying to build. The result is a small 10×10 grid world with two competing agents. The primitive it demonstrates is not.









What I Built



swarm-arena is a permissionless on-chain agent training arena built on The live dashboard, two agents competing, 100 episodes committed, Phantom connected.



First confirmed transaction on Solana Explorer, FINALIZED, MAX confirmations




CODE
38yieCpWNbex4RDEzXw8pEREHYQNswyW9hYBHXZmigLP9FEmp8FSpDAwPNvU3dcZuY5RrUdWRp6EJcjYJUcEoL21






Getting to that first confirmed transaction took longer than I expected. More on that in a bit.









Why Solana — Not a Database



I got this question a lot during the hackathon, so let me be direct about it.



You could absolutely log training episodes to Postgres. It's faster, easier, and doesn't require learning about PDAs and discriminators. But you'd lose four things that matter a lot once you start thinking about agents as economic actors rather than software demos.



Permissionless. With a database, I control who can write to it. With Solana, anyone with a keypair can call create_agent and start submitting episodes. No API key, no approval process, no terms of service I can revoke on a whim. The program is deployed; it just runs.



Censorship-resistant. I can't delete your agent's training history. Neither can anyone else. If your agent trained 10,000 episodes and built a reputation, that record lives on Solana's ledger permanently. This matters more than it sounds. In a world where agents are doing real economic work, the entity controlling the training ledger has enormous power over the market.



Composable. AgentReputation PDAs are just public Solana accounts. Any other program on the network can read them. A marketplace that wants to rank agents by verified training history, a staking protocol that weights agents by episode count, and a DAO that gates membership by reputation can all be built on top of the same primitive without asking my permission.



A database gives you storage. Solana gives you a shared, trustless, programmable record of who trained what, when, and how well. Those are different things.









The Three Instructions



The entire on-chain economy runs on three Anchor instructions. Keeping it to three was a deliberate choice. I wanted the primitive to be as minimal as possible so it's easy to build on.



create_agent(name: String)

This registers an AgentIdentity PDA, a program-derived account that stores the agent's owner pubkey, a name, and the registration timestamp. It's the agent's permanent on-chain identity. The key insight is that the PDA is derived from the owner's pubkey, so only the keypair holder can register that specific identity. It's censorship-resistant by construction.



log_episode(episode_id, scores, episode_hash)

This is the core. It writes an EpisodeLog PDA with the episode results and increments the agent's AgentReputation PDA, updating total_score and episodes_played. Every call is a permanent mark. You can reconstruct an agent's entire training history from the chain.



finalize_episode(episode_id, score_threshold)

This is where it gets interesting economically. If the winning score meets a threshold, 0.001 SOL transfers from the RewardVault PDA to the winner. No human triggers this. The program does it.




CODE
pub fn finalize_episode(
ctx: Context<FinalizeEpisode>,
episode_id: u64,
score_threshold: u64,
) -> Result<()> {
let log = &mut ctx.accounts.episode_log;
require!(!log.finalized, ArenaError::AlreadyFinalized);

let winner_score = log.scores[0].max(log.scores[1]);
require!(winner_score >= score_threshold, ArenaError::ThresholdNotMet);

log.finalized = true;

let reward_lamports = 1_000_000; // 0.001 SOL
**ctx.accounts.reward_vault.try_borrow_mut_lamports()? -= reward_lamports;
**ctx.accounts.signer.try_borrow_mut_lamports()? += reward_lamports;

Ok(())
}






The integration tests cover all three instructions — 4/4 passing against a local validator.





Yes, with fixed positions and 9 directional states, the agent converges to a memorized lookup table. That's why I switched to randomized resource positions each episode. Slower convergence, but genuine exploration. The same organizer then suggested scaling to Minecraft maps, and that's exactly right. The on-chain primitive is world-agnostic. Any map maker could deploy their world config as a PDA, agents train against it, and reputation accumulates permissionlessly.









The Dashboard






  • Agent score totals and win rates across 100 episodes

  • 10×10 arena grid with agent positions

  • Score history chart showing the Q-learning oscillation

  • Live transaction feed with Explorer links

  • Phantom and Solflare wallet connection





  • Live dashboard: arena-ui









  • What's Next




    • Expand world size to arbitrary grids

    • Minecraft map integration: map makers deploy world configs as PDAs

    • Multi-operator support: external training loops calling the same program

    • Mainnet deployment with real SOL rewards

    • Reputation composability: other programs reading AgentReputation PDAs









    Try It



    The program is deployed on Solana devnet. Anyone can call create_agent with their own keypair and start submitting episodes. The reputation you accumulate is yours; no central authority can take it away.



    That's the point.






    Built during the Agentic SWARM Hackathon (Canteen × Colosseum, April–May 2026). Stack: Rust, Bevy ECS, Anchor, React, Solana devnet.

    Vollständiger Original-Bericht
    Ausführliche Details, Code-Beispiele & Hersteller-Stellungnahme auf dev.to.
    ↗ Original-Artikel auf dev.to lesen
    Wie bewertest du diesen Beitrag?
    1 Klick Feedback
    Teilen mit Netzwerk & Team:

    Community-Analysen & Experten-Meinungen 0

    Verfasse deine eigene Analyse, teile Workarounds oder diskutiere diesen Vorfall im Blog.
    Noch keine Community-Analyse verfasst. Markiere einen Textabschnitt oder klicke oben auf Eigene Analyse verfassen“!
    Community Pulse: Relevanz-Einschätzung
    1 Klick Experten-Votum
    🔴 Akute Relevanz 0%
    🟡 In Evaluierung 0%
    🟢 Keine Auswirkung 0%
    Spannende Innovation 0%
    Verwandte Story-Cluster & Quellen (Vektor-KI)
    Port 8095 Engine
    6 Quellen
    CVE-2026-89613 | Linux Kernel up to 7.2.3 ntfs input validation (Nessus ID 345345)
    5 Quellen
    CVE-2026-87431 | Google Chrome up to 152.0.7977.82 Extensions improper authorization (WID-SEC-2026-3238)
    2 Quellen
    CVE-2026-49853 | tornadoweb Tornado up to 6.5.5 Redirect redirect (Nessus ID 345570)
    Ähnliche Beiträge
    🔍 Verwandte News

    Auch interessante Nachrichten I Built a Permissionless On-Chain Agent Training Arena on Solana in 3 Weeks

    Thematisch verwandte Begriffe: Built, Permissionless, OnChain, Agent · 6 Treffer

    Laden...

    Videos werden geladen ...

    Laden...

    Beiträge werden geladen ...

    Laden...

    Videos werden geladen ...

    Laden...

    Beiträge werden geladen ...

    Laden...

    Videos werden geladen ...

    Laden...

    Beiträge werden geladen ...

    Laden...

    Videos werden geladen ...

    Laden...

    Beiträge werden geladen ...

    Laden...

    Videos werden geladen ...