⚠️ Malware / Trojaner / VirenAndroid-Malware blockiert Google Play per VPN-Trick(24.08.2026 um 13:00 Uhr)
🪟 Windows TippsWindows 11: Falsche Defender-Warnungen und kaputte Mauszeiger(31.08.2026 um 11:58 Uhr)
🕵️ SicherheitslückenDropbox-Hack: Tausende Konten kompromittiert(02.09.2026 um 11:02 Uhr)
⚠️ Malware / Trojaner / VirenAtombomben-Frage trickst KI-Malware-Scanner aus(02.09.2026 um 12:46 Uhr)
🪟 Windows TippsMehr Sicherheit in Windows 11(03.09.2026 um 11:51 Uhr)
⚠️ Malware / Trojaner / VirenNeue Android-Malware schreit Sie an, wenn Sie nicht zahlen(11.09.2026 um 10:33 Uhr)
🐧 Linux TippsMehrere Probleme in freerdp2 (Fedora)(11.09.2026 um 23:24 Uhr)
🐧 Linux TippsMehrere Probleme in kamailio (Debian)(11.09.2026 um 23:28 Uhr)
🐧 Linux TippsAusführen beliebiger Kommandos in dokuwiki (Fedora)(11.09.2026 um 23:28 Uhr)
🐧 Linux TippsAusführen beliebiger Kommandos in python-asteval (Fedora)(11.09.2026 um 23:28 Uhr)
⚠️ Malware / Trojaner / VirenAndroid-Malware blockiert Google Play per VPN-Trick(24.08.2026 um 13:00 Uhr)
🪟 Windows TippsWindows 11: Falsche Defender-Warnungen und kaputte Mauszeiger(31.08.2026 um 11:58 Uhr)
🕵️ SicherheitslückenDropbox-Hack: Tausende Konten kompromittiert(02.09.2026 um 11:02 Uhr)
⚠️ Malware / Trojaner / VirenAtombomben-Frage trickst KI-Malware-Scanner aus(02.09.2026 um 12:46 Uhr)
🪟 Windows TippsMehr Sicherheit in Windows 11(03.09.2026 um 11:51 Uhr)
⚠️ Malware / Trojaner / VirenNeue Android-Malware schreit Sie an, wenn Sie nicht zahlen(11.09.2026 um 10:33 Uhr)
🐧 Linux TippsMehrere Probleme in freerdp2 (Fedora)(11.09.2026 um 23:24 Uhr)
🐧 Linux TippsMehrere Probleme in kamailio (Debian)(11.09.2026 um 23:28 Uhr)
🐧 Linux TippsAusführen beliebiger Kommandos in dokuwiki (Fedora)(11.09.2026 um 23:28 Uhr)
🐧 Linux TippsAusführen beliebiger Kommandos in python-asteval (Fedora)(11.09.2026 um 23:28 Uhr)

🔧 Programmierung 🕛 vor 3 Monaten 7 Min Lesezeit
0

NVIDIA and Apple Solved the Hardware. Here's What's Left to Build.

↗ Quelle (dev.to)
🗣️ Stimme:
📑 Inhaltsübersicht

After GTC 2026, one thing is basically settled: the hardware layer for on-device AI is no longer the bottleneck.



NVIDIA's RTX Spark packs Blackwell GPU + Grace CPU + 128GB unified memory into a desktop form factor. Apple's M-series chips with unified memory architecture and efficiency-first design let 4B and even 7B parameter models run smoothly on a MacBook. Two different approaches, same destination: consumer hardware now has the compute foundation for running on-device AI agents.



Chip vendors have done their part. The next question is: how many layers are still missing between "chip can run an AI model" and "an on-device agent can actually complete useful tasks"?



This post maps out the full technology stack for on-device AI agents, examining each layer's maturity, identifying gaps, and tracking what the open-source community has built so far.






Layer 1: Silicon (Ready)



On-device AI inference has different chip requirements than traditional compute workloads. The core bottleneck isn't peak FLOPS — it's memory bandwidth and unified memory capacity. LLM inference needs model weights fully loaded into memory, with high-frequency data movement between weight matrices and activations during computation. If memory bandwidth can't keep up, raw compute power just sits idle waiting for data.



Three main silicon paths exist today:





  • NVIDIA N1X: Blackwell GPU + Grace CPU heterogeneous architecture, 128GB unified memory, petaflop-class compute, targeting desktop workstations


  • Apple M-series (M4/M5): Unified memory architecture with GPU and CPU sharing memory, optimized memory bandwidth, configurations from 32GB to 192GB


  • Qualcomm Snapdragon X: Targeting laptops and mobile, NPU-accelerated inference, relatively limited memory configurations



Different emphases, but one common takeaway: 2026 consumer silicon can run 4B+ parameter models for real-time inference. This layer is ready.






Layer 2: Inference Frameworks (Mature)



With silicon in place, efficient inference frameworks are needed to actually run models. This layer solves the problem of mapping deep learning models efficiently onto specific chip compute units.



Apple ecosystem: MLX is the most mature inference framework on Apple Silicon. Native support for weight quantization (W8A16, W4A16), deep Metal GPU optimization, active community.



NVIDIA ecosystem: TensorRT-LLM is the corresponding solution, optimized for CUDA and Tensor Cores, with specific adaptations for Blackwell architecture on RTX Spark.



Cross-platform: ONNX Runtime for multi-platform deployment, llama.cpp taking the minimalist approach running on diverse hardware.



This layer is mature enough. Developers don't need to write inference kernels from scratch — pick a framework and your model runs.






Layer 3: Quantization Acceleration (Catching Up)



Inference frameworks make models "runnable." The quantization acceleration layer makes them "fast."



The computational bottleneck in LLM inference is matrix multiplication. Model weights are typically stored in FP16 or BF16, but edge chips have dedicated hardware acceleration units for low-precision compute. Quantizing weights and activations to INT8 or INT4 significantly improves inference speed and reduces memory footprint.



SDK fills this gap. Built on top of MLX, Cider implements W8A8 and W4A8 activation quantization modes, quantizing both weights and activations to INT8 for direct INT8 TensorOps matrix multiplication. Measured performance:




  • On Apple M5 Pro, W8A8 per-channel quantization achieves up to 1.8x prefill speedup over W8A16 baseline

  • Compared to MLX native W4A16, prefill speedup ranges from 1.4x to 2.2x

  • Compatible with all MLX models, not limited to any specific project



Cider uses conditional compilation: M5+ chips get the full C++ extension and Metal kernels built; M4 and below install as a pure-Python package for compatibility fallback. Different hardware, same install command, but acceleration only kicks in on M5+.



This layer is in the "catching up" phase. Weight quantization is standard. Activation quantization is becoming mainstream. Finer-grained strategies (per-group, per-token) are still evolving.






Layer 4: Models (Usable in Vertical Domains)



The first three layers are infrastructure. Layer 4 is where the model directly faces the task. The core challenge for on-device models: parameter count is constrained by device memory, but task complexity doesn't decrease just because you're running locally.



The generic approach distills or prunes cloud-scale models down to on-device size, but this typically comes with noticeable capability degradation.



A more effective path is domain-specific optimization. Through targeted training on specific task types (GUI operations, web navigation, code generation), small models can match or exceed large models on their target domains.





Benchmark data (72B evaluation model):




  • OSWorld: 58.2% accuracy, #1 among specialized models, leading second-place opencua-72b (45.0%) by 13.2 percentage points

  • WebRetriever Protocol I: 41.7 NavEval, ahead of Gemini 2.5 Pro at 40.9 and Claude 4.5 at 31.3



Note: these results are from the 72B evaluation model. The actual on-device deployment uses the 4B version (Mano-CUA-4B-Thinking-1.1), achieving roughly 80 tokens/s decode speed on M5 Pro with 64GB RAM. With Cider's W8A8 quantization, prefill gets an additional ~12.7% speedup over the W8A16 baseline.



This layer's status: general capability still has a gap, but in vertical domains like GUI operations and web navigation, on-device specialized models are production-ready.






Layer 5: Agent Orchestration (Early Engineering)



A model that can understand instructions and operate interfaces still needs an orchestration layer to manage task decomposition, tool invocation, error recovery, and state tracking to complete full workflows.



The challenge here: on-device agents can't rely on massive cloud compute for complex planning and backtracking. All decisions must happen within local resource constraints.








What This Means for Developers



If you're building in the on-device AI space, this is a window worth paying attention to. The silicon and framework layers are mature. Quantization and model layers are iterating rapidly. Getting involved now puts you in the critical phase where the ecosystem moves from "works" to "works well."



Your specific stack choices depend on your use case:





  • Quick validation of on-device GUI agent capabilities: Use Mano-P's cloud mode (via mano.mininglamp.com) to get started, then switch to local mode


  • Inference acceleration optimization on Apple Silicon: Cider's INT8 TensorOps implementation is a useful reference


  • Building end-to-end autonomous task pipelines: Mano-AFK's architecture (separate builder agent + adversary reviewer agent) is worth studying



All projects are open-source under the Mininglamp-AI GitHub organization. Mano-P is Apache 2.0 licensed, installable via brew tap Mininglamp-AI/tap && brew install mano-cua. If you find the work useful, a GitHub star goes a long way.

Vollständiger Original-Bericht
Ausführliche Details, Code-Beispiele & Hersteller-Stellungnahme auf dev.to.
↗ Original-Artikel auf dev.to lesen
Wie bewertest du diesen Beitrag?
1 Klick Feedback
Teilen mit Netzwerk & Team:

Community-Analysen & Experten-Meinungen 0

Verfasse deine eigene Analyse, teile Workarounds oder diskutiere diesen Vorfall im Blog.
Noch keine Community-Analyse verfasst. Markiere einen Textabschnitt oder klicke oben auf Eigene Analyse verfassen“!
Community Pulse: Relevanz-Einschätzung
1 Klick Experten-Votum
🔴 Akute Relevanz 0%
🟡 In Evaluierung 0%
🟢 Keine Auswirkung 0%
Spannende Innovation 0%
Verwandte Story-Cluster & Quellen (Vektor-KI)
Port 8095 Engine
1 Quelle
Stealing AI Reasoning Traces
1 Quelle
AIs as Modern Genies
1 Quelle
Bitcoin: KI-Hacker räumen Millionen ab! Wird Künstliche Intelligenz zum Problem? - ftd.de
Ähnliche Beiträge
🔍 Verwandte News

Auch interessante Nachrichten NVIDIA and Apple Solved the Hardware. Here's What's Left to Build.

Thematisch verwandte Begriffe: NVIDIA, Apple, Solved, Hardware · 6 Treffer

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...