🤖💻 AI Daily Digest — July 12, 2026
Another packed week in AI. OpenAI ended its 12-day restricted preview and opened GPT-5.6 to the world — three models, a new durability concept, and the ChatGPT Work + Codex integration that signals where the company is headed. Meta's Muse Spark 1.1 landed with enough firepower to pull Mark Zuckerberg back to X after three years of silence. And NVIDIA and Hugging Face took a big swing at open-source robotics.
Let's dig in.
1. OpenAI GPT-5.6 Goes Public — Sol, Terra, Luna Redefine the Tier System
OpenAI officially released the GPT-5.6 series on July 9, ending 12 days of restricted government preview. The three-model family — Sol, Terra, and Luna — introduces a new "durable capability tier" concept: the names identify capability levels, not versions, meaning Sol can be upgraded to a future GPT-5.7 while keeping its tier identity.
Sol, the flagship, sets new state-of-the-art across coding (80 on the Artificial Analysis Coding Agent Index, beating Claude Fable 5 by 2.8 points), cybersecurity (73.5% on ExploitBench vs GPT-5.5's 47.9%), and knowledge work. It runs in three effort modes: default for cost efficiency, max for extended reasoning, and ultra which coordinates 4 parallel agents by default (scalable to 16). Pricing runs $5/$30 per million input/output tokens for Sol, $2.50/$15 for Terra, and $1/$6 for Luna.
Alongside the model launch, OpenAI merged Codex into the ChatGPT desktop app and introduced ChatGPT Work, a unified interface for chat, coding, and long-running agent tasks. A new Programmatic Tool Calling feature in the Responses API lets GPT-5.6 write and run lightweight programs that coordinate tools inline. According to internal benchmarks, GPT-5.6 Sol improved the RSI Index by 16.2 points over GPT-5.5 on AI research acceleration tasks.
— OpenAI · ChatGPT Blog
🔗
2. OpenAI GPT-Live — Real-Time Voice That Actually Listens and Speaks Simultaneously
The same week, OpenAI launched GPT-Live, a full-duplex voice model that listens and speaks at the same time. Two versions — GPT-Live-1 and GPT-Live-1 mini — started rolling out globally on July 8.
Previous voice systems either chained three models together (cascaded) or worked in rigid turn-based mode where the model waited for silence before responding. GPT-Live's full-duplex architecture processes input continuously while generating output, making interaction decisions many times per second — whether to speak, listen, pause, interrupt, or invoke a tool. It handles backchannel cues ("mhmm", "yeah"), stays quiet when you need a moment, and can perform real-time simultaneous translation.
When a question requires deeper reasoning or search, GPT-Live delegates to GPT-5.5 behind the scenes and brings results back into the conversation without breaking flow. In head-to-head evaluations, both models are strongly preferred over Advanced Voice Mode for pleasantness, turn-taking, and natural flow. GPT-Live-1 substantially outperforms Advanced Voice Mode on GPQA (expert-level science reasoning) and BrowseComp (agentic web search). A demo showed it translating live between languages with no perceptible delay.
— OpenAI
🔗 ·
4. Microsoft Swaps In-House MAI Models Into Excel and Outlook
Microsoft has quietly started replacing third-party AI models — including OpenAI's and Anthropic's — with its in-house MAI series in core Office products, Bloomberg reported on July 8.
Excel and Outlook now process tens of thousands of weekly AI prompts entirely on MAI models, a deployment scale not previously disclosed. While the swap covers only a fraction of Microsoft's total AI workload, it marks a strategic inflection point: Microsoft is no longer willing to pay premium pricing to OpenAI and Anthropic at scale. Mustafa Suleyman's AI team is building toward full model independence, with the MAI series designed to handle Copilot's massive token consumption at a fraction of the cost. The current OpenAI partnership still provides discounted access, but those terms are narrowing.
— Bloomberg · Peng
🔗 ·
7. Mistral Leanstral 1.5: Open-Source Formal Verification That Finds Real Bugs
Mistral AI released Leanstral 1.5 under Apache 2.0 on July 2 — a 119-billion-parameter sparse MoE model specialized for theorem proving and code verification in Lean 4.
The numbers are striking: 100% on miniF2F (both validation and test), 587 out of 672 PutnamBench problems solved, and new state-of-the-art on FATE-H (87%) and FATE-X (34%) algebra verification benchmarks. At roughly $4 per problem on PutnamBench, it's far below the highest-compute comparison systems.
But the real story is practical impact: Mistral used a pipeline translating Rust into Lean, generating candidate correctness properties, and attempting to prove or disprove them. Across 57 open-source repositories, Leanstral 1.5 identified 11 genuine bugs, five of which were previously unreported. The model activates only ~6 billion parameters per token (of 119B total), making it deployable at a fraction of its full compute cost. It supports a 256,000-token context window and is available free through Mistral's Labs tier and Vibe agent environment.
— Mistral AI
🔗 · Hugging Face Model
Next digest: July 13, 2026. Follow KD Agentic for daily AI coverage.
SOCIAL SHARE CARD GENERATOR