Zum Hauptinhalt springen
Echtzeit-Radar & Feeds
Alle RSS Feeds ➔
👥 Community & Social
Linux Tipps & HardeningSecurity: Ausführen beliebiger Kommandos in perl-Dancer2 (Fedora)(29.09.2026 um 07:43 Uhr)
•
Linux Tipps & HardeningSecurity: Denial of Service in perl-HTML-FormFu (Fedora)(29.09.2026 um 07:46 Uhr)
••
Linux Tipps & HardeningSecurity: Zwei Probleme in NetworkManager-l2tp (Fedora)(29.09.2026 um 07:46 Uhr)
••
Sicherheitslücken (CVE)CVE-2026-77144 | TYPO3 Events 2 Plugin up to 10.2.11 permission(29.09.2026 um 06:21 Uhr)
•
Sicherheitslücken (CVE)CVE-2026-21753 | HCL Hive 1.0 unmaintained third party components(29.09.2026 um 06:21 Uhr)
••
Sicherheitslücken (CVE)CVE-2026-75038 | ilya-zlobintsev LACT up to 0.10.0 symlink(29.09.2026 um 06:21 Uhr)
••
Linux Tipps & HardeningSecurity: Ausführen beliebiger Kommandos in perl-Dancer2 (Fedora)(29.09.2026 um 07:43 Uhr)
•
Linux Tipps & HardeningSecurity: Denial of Service in perl-HTML-FormFu (Fedora)(29.09.2026 um 07:46 Uhr)
••
Linux Tipps & HardeningSecurity: Zwei Probleme in NetworkManager-l2tp (Fedora)(29.09.2026 um 07:46 Uhr)
••
Sicherheitslücken (CVE)CVE-2026-77144 | TYPO3 Events 2 Plugin up to 10.2.11 permission(29.09.2026 um 06:21 Uhr)
•
Sicherheitslücken (CVE)CVE-2026-21753 | HCL Hive 1.0 unmaintained third party components(29.09.2026 um 06:21 Uhr)
••
Sicherheitslücken (CVE)CVE-2026-75038 | ilya-zlobintsev LACT up to 0.10.0 symlink(29.09.2026 um 06:21 Uhr)
••
Intelligence View
⚡ tsecurity.de Intelligence

Why sharding is essential to fine-tuning LLMs

This article was originally published on IBM Developer. Training and fine-tuning large language models (LLMs) is becoming a central requirement for modern AI applications. As these models grow in size—from billions to hundreds of billions …

0
↗ Quelle (dev.to)
Reagiere als Erste:r — dein Feedback zählt!

This article was originally published on IBM Developer.



Training and fine-tuning large language models (LLMs) is becoming a central requirement for modern AI applications. As these models grow in size—from billions to hundreds of billions of parameters—the demands on computational resources have increased dramatically. Fine-tuning such models on a single GPU is no longer realistic due to memory limitations and training inefficiencies.



Sharding is the process of splitting a model’s data or components across multiple devices—such as GPUs or nodes—so that the training workload is distributed. By dividing the model’s parameters, gradients, and optimizer states into smaller “shards,” each device only needs to manage a fraction of the total, making it possible to train models that would not otherwise fit in memory. Sharding also enables parallel training, which speeds up the process and improves scalability.



In this article, we explore the importance of sharding for scalable LLM fine-tuning, describe various sharding strategies, and provide practical guidance based on industry-standard tools.






Why training and fine-tuning LLMs require sharding



Training large language models (LLMs) involves handling substantial amounts of data and computation during each pass through the network. These passes are generally referred to as:





  • Forward pass: When data flows through the model to generate predictions.


  • Backward pass: When the model computes how wrong the predictions were (loss) and adjusts internal weights accordingly through backpropagation.



Each training iteration requires tracking and updating several core components:





  • Model parameters: These are the learnable weights of the neural network that determine the model’s behaviour. They are updated during training to minimize prediction errors.


  • Gradients: These represent the rate of change of the loss with respect to each model parameter. Gradients are computed during the backward pass and guide how the model updates its parameters.


  • Optimizer states: These are internal values maintained by optimization algorithms like Adam or SGD. They help fine-tune how each parameter gets updated based on the gradient and previous updates.



While inference can be managed on a single GPU using techniques like offloading or quantization, training requires all three of these components to reside in GPU memory simultaneously. This can triple the memory requirement compared to inference. Without sharding, even relatively modest models (7B–13B parameters) can exceed the capabilities of high-end GPUs.



Moreover, sharding enables:




  • Larger batch sizes, improving convergence and model generalization.

  • Distributed compute workloads, reducing training time.

  • Better scalability across infrastructure.



Continue reading on IBM Developer to see a DeepSpeed ZeRO example of scalable fine-tuning...

2. Cyber Threat Intelligence & Forensik

CTI Threat Relationship Graph2 Knoten / 1 Relationen
CVE / Incident Software MITRE ATT&CK CWE Weakness IoC
Ähnliche Beiträge
🔍 Verwandte News

Auch interessante Nachrichten Why sharding is essential to fine-tuning LLMs

Thematisch verwandte Begriffe: sharding, essential, finetuning, LLMs · 6 Treffer

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

💬 Kommentare werden geladen…
Zum Aktualisieren ziehen
ZERO-DAY CVE-2026-102367 | mall4j through 4.0 contains an insufficient session expiration vulnerab…
Advisory →
tsecurity.de Icon
Offline-Lesen, Eilmeldungen & 0ms Ladezeit

Installiere tsecurity.de direkt auf deinen Home-Bildschirm für das ultimative Vollbild-Magazinerlebnis ohne Browser-Leisten.

Nächster Beitrag