🪟 Windows TippsHow to Enable Windows 11 Screen Savers(07.09.2026 um 12:41 Uhr)
🪟 Windows TippsMicrosoft Phone Link Not Showing Messages on Windows 11? Fix It(09.09.2026 um 07:52 Uhr)
⚠️ Malware / Trojaner / VirenPost-DEF CON phishing campaign delivered AMOS and NetSupport malware(24.08.2026 um 09:42 Uhr)
💾 IT Security ToolsHow to Use BloodHound Active Directory Setup Attack Path Analysis(10.09.2026 um 14:35 Uhr)
🕵️ SicherheitslückenCompliance Alert: EU Cyber Resilience Act 24-Hour Reporting Enforced(11.09.2026 um 06:25 Uhr)
🕵️ SicherheitslückenAWS IAM Privilege Escalation: Cheat Sheet And Defense(11.09.2026 um 07:43 Uhr)
🕵️ SicherheitslückenArista warns customers ahead of next week’s security disclosures(02.09.2026 um 23:34 Uhr)
🕵️ SicherheitslückenKARR Security vulnerability(02.09.2026 um 03:15 Uhr)
🪟 Windows TippsHow to Enable Windows 11 Screen Savers(07.09.2026 um 12:41 Uhr)
🪟 Windows TippsMicrosoft Phone Link Not Showing Messages on Windows 11? Fix It(09.09.2026 um 07:52 Uhr)
⚠️ Malware / Trojaner / VirenPost-DEF CON phishing campaign delivered AMOS and NetSupport malware(24.08.2026 um 09:42 Uhr)
💾 IT Security ToolsHow to Use BloodHound Active Directory Setup Attack Path Analysis(10.09.2026 um 14:35 Uhr)
🕵️ SicherheitslückenCompliance Alert: EU Cyber Resilience Act 24-Hour Reporting Enforced(11.09.2026 um 06:25 Uhr)
🕵️ SicherheitslückenAWS IAM Privilege Escalation: Cheat Sheet And Defense(11.09.2026 um 07:43 Uhr)
🕵️ SicherheitslückenArista warns customers ahead of next week’s security disclosures(02.09.2026 um 23:34 Uhr)
🕵️ SicherheitslückenKARR Security vulnerability(02.09.2026 um 03:15 Uhr)

🔧 Programmierung 🕛 vor 2 Jahren 3 Min Lesezeit
0

How to Serve LLM Completions in Production

↗ Quelle (dev.to)
🗣️ Stimme:
📑 Inhaltsübersicht




Preparations



To start, you need to compile for instructions.



The server is compiled alongside other targets by default.



Once you have the server running, we can continue. We will use PHP Resonance framework.






Troubleshooting






Obtaining Open-Source LLM



I recommend starting either with . You need to download the pretrained weights and convert them into GGUF format before they can be used with supports CPU-only setups, so you don't have to do any additional configuration. It will be slow, but you will still have tokens generated.






Running With a Low VRAM Memory



You can try quantization if you don't have enough VRAM on your GPU to run a specific model. That lowers the response quality and the memory the model needs to use. Llama.cpp has a utility to quantize models:




CODE
$ ./quantize ./models/7B/ggml-model-f16.gguf ./models/7B/ggml-model-q4_0.gguf q4_0






10GB of VRAM is enough to run most quantized models.






Starting llama.cpp Server



While writing this tutorial, I had a server started with a command:




CODE
$ ./server 
--model ~/llama-2-7b-chat/ggml-model-q4_0.gguf
--n-gpu-layers 200000
--ctx-size 2048
--parallel 8
--cont-batching
--mlock
--port 8081






cont-batching parameter is essential, because it enables continuous batching, which is an optimization technique that allows parallel request.



Without it, even with multiple parallel slots, the server could answer to only one request at a time. cont-batching allows the server to respond to multiple completion requests in parallel.






Configuring Resonance



All you need to do is add a configuration section that specifies the llama.cpp server location:




CODE
[llamacpp]
host = 127.0.0.1
port = 8081









Testing



Resonance has built-in commands that connect to llama.cpp and issue requests.



You can send a sample prompt through llamacpp:completion:




CODE
$ php ./bin/resonance.php llamacpp:completion "How to write a 'Hello, world' in PHP?"
To write a "Hello, world" in PHP, you can use the following code:

<?php
echo "Hello, world!";
?>

This will produce a simple "Hello, world!" message when executed.









Programmatic Use



In your class, you need to use server and connect to it with Resonance.

Vollständiger Original-Bericht
Ausführliche Details, Code-Beispiele & Hersteller-Stellungnahme auf dev.to.
↗ Original-Artikel auf dev.to lesen
Wie bewertest du diesen Beitrag?
1 Klick Feedback
Teilen mit Netzwerk & Team:

Community-Analysen & Experten-Meinungen 0

Verfasse deine eigene Analyse, teile Workarounds oder diskutiere diesen Vorfall im Blog.
Noch keine Community-Analyse verfasst. Markiere einen Textabschnitt oder klicke oben auf Eigene Analyse verfassen“!
Community Pulse: Relevanz-Einschätzung
1 Klick Experten-Votum
🔴 Akute Relevanz 0%
🟡 In Evaluierung 0%
🟢 Keine Auswirkung 0%
Spannende Innovation 0%
Verwandte Story-Cluster & Quellen (Vektor-KI)
Port 8095 Engine
1 Quelle
Kritische „OVERPASS“-Lücke bedroht den SAP-Kernel - it-daily.net
1 Quelle
ChatGPT: Versteckter Prompt konnte Gmail-Daten abgreifen - it-daily.net
1 Quelle
GuardBreaker: Derailing AI-assisted malware analysis with a code comment
Ähnliche Beiträge
🔍 Verwandte News

Auch interessante Nachrichten How to Serve LLM Completions in Production

Thematisch verwandte Begriffe: Serve, Completions, Production · 6 Treffer

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...