Zum Hauptinhalt springen
tsecurity.de LIVE
Echtzeit-Radar & Feeds
Alle RSS Feeds
👥 Community & Social
IT Security NachrichtenKI-Agenten hebeln klassisches IT-Asset-Management aus(24.09.2026 um 07:14 Uhr)
IT Security NachrichtenMicrosoft erneuert Surface Pro und Laptop mit Snapdragon X2 Plus(24.09.2026 um 07:42 Uhr)
IT Security DownloadsGitHub Release: cline/cline vsdk/shared/v0.0.86 (24.09.2026)(24.09.2026 um 07:43 Uhr)
IT Security DownloadsGitHub Release: cline/cline vsdk/llms/v0.0.86 (24.09.2026)(24.09.2026 um 07:43 Uhr)
IT Security DownloadsGitHub Release: cline/cline vsdk/agents/v0.0.86 (24.09.2026)(24.09.2026 um 07:43 Uhr)
IT Security DownloadsGitHub Release: cline/cline vsdk/core/v0.0.86 (24.09.2026)(24.09.2026 um 07:43 Uhr)
IT Security DownloadsGitHub Release: cline/cline vcli-v3.0.65 (24.09.2026)(24.09.2026 um 07:54 Uhr)
IT Security NachrichtenKI-Agenten hebeln klassisches IT-Asset-Management aus(24.09.2026 um 07:14 Uhr)
IT Security NachrichtenMicrosoft erneuert Surface Pro und Laptop mit Snapdragon X2 Plus(24.09.2026 um 07:42 Uhr)
IT Security DownloadsGitHub Release: cline/cline vsdk/shared/v0.0.86 (24.09.2026)(24.09.2026 um 07:43 Uhr)
IT Security DownloadsGitHub Release: cline/cline vsdk/llms/v0.0.86 (24.09.2026)(24.09.2026 um 07:43 Uhr)
IT Security DownloadsGitHub Release: cline/cline vsdk/agents/v0.0.86 (24.09.2026)(24.09.2026 um 07:43 Uhr)
IT Security DownloadsGitHub Release: cline/cline vsdk/core/v0.0.86 (24.09.2026)(24.09.2026 um 07:43 Uhr)
IT Security DownloadsGitHub Release: cline/cline vcli-v3.0.65 (24.09.2026)(24.09.2026 um 07:54 Uhr)
Intelligence View
⚡ tsecurity.de Intelligence

Design and Implementation of a Slurm-Based HPC Cluster

The Problem Managing a growing fleet of GPU and HPC servers one-by-one doesn't scale - here's how we fixed it. Our department has several computing resources, including GPU servers and HPC servers. Previously, these machines were managed…

0
↗ Quelle (dev.to)
Reagiere als Erste:r — dein Feedback zählt!

The Problem

Managing a growing fleet of GPU and HPC servers one-by-one doesn't scale - here's how we fixed it.



Our department has several computing resources, including GPU servers and HPC servers. Previously, these machines were managed and accessed individually. As the number of servers and users grew, so did the inefficiencies.



From an administrator's perspective, it was difficult to monitor resource usage and ensure fair utilization across machines. Some servers would become heavily loaded while others sat idle, with no automatic balancing in place.



From a user's perspective, researchers and students had to manually decide which server to use - without any visibility into current availability or load.



To address these challenges, we built a Slurm-based HPC cluster. Users now submit jobs with their resource requirements, and the scheduler automatically selects an appropriate compute node. This simplifies resource management and allows the department's computing infrastructure to be utilized far more efficiently.



What Is Slurm?



Slurm (Simple Linux Utility for Resource Management) is an open-source workload manager widely used in HPC environments. It handles job queuing, scheduling, and resource allocation across a cluster of machines — letting users focus on their work rather than infrastructure logistics.



Architecture



SLURM Architecture



The cluster consists of three main components:





  • Login node — the entry point where users connect and submit jobs


  • Controller node (CERF) — runs slurmctld, the central Slurm daemon responsible for scheduling decisions


  • Compute nodeskepler (GPU node) and aiken (HPC node), each running slurmd to receive and execute assigned jobs



Users submit jobs to the controller, which schedules and dispatches them to the appropriate compute node based on resource availability and job requirements.



Communication between Slurm components is secured using Munge authentication, which establishes trust between cluster nodes and ensures that scheduling operations are performed securely. User authentication is handled separately through the standard Linux user management system.



A Slurm partition groups the available compute resources, and the controller maintains real-time information about each node's state — allowing it to make informed scheduling decisions.



Submitting Jobs



Once the cluster was operational, users could submit workloads through standard Slurm commands. Instead of SSH-ing directly into a compute node, users interact with the cluster through the login node using commands like:





  • srun — run a command interactively on an allocated node


  • sbatch — submit a batch script to be executed asynchronously


  • squeue — view the current job queue and job statuses


  • scancel — cancel a running or pending job



The request is sent to the Slurm controller, which identifies a suitable compute node and dispatches the job for execution.



The figure below shows a simple srun command being submitted from the login node, scheduled by the controller, and executed on the kepler GPU node — with the output returned directly to the user's terminal.



Example

CTI Threat Relationship Graph2 Knoten / 1 Relationen
CVE / Incident Software MITRE ATT&CK CWE Weakness IoC
SOC Incident Playbook: Remote Code Execution (RCE) Defense
title: Detect Exploitation - Design and Implementation of a Slurm-Based HPC Cluster
id: e56c07eb-9cd3-47c3-8dc2-319c136bf589
status: experimental
description: Automatisch generierte SIEM-Erkennungsregel basierend auf CTI Intelligence
references:
  - https://tsecurity.de/
author: iShareStuff CTI Automated Detection Engine
date: 2026-09-24
logsource:
  category: network_connection
  product: any
detection:
  selection:
      CommandLine|contains:
        - 'exploit'
  condition: selection
falsepositives:
  - Legitime administrative Zugriffe oder Penetrationstests
level: high
tags:
  - attack.initial_access
rule CTI_Threat_Indicator {
    meta:
        author = "iShareStuff CTI Automated Detection Engine"
        date = "2026-09-24"
        description = "YARA Signature for "
    strings:
        $str = "Design and Implementation of a" ascii wide
    condition:
        any of them
}
tsecurity.de Cognitive Threat RAG
Fokus-Vektor:

Kognitive Analyse für identifizierte Bedrohung: Erhöhte Bedrohungslage im Bereich Design and Implementation of a Slurm-Bas.... Basierend auf 368k Vektor-Korrelationen werden sofortige Isolationsmaßnahmen für betroffene Endpunkte empfohlen.

🛡️ Angriffsfläche & Exposure

Netzwerk/Remote-Zugriff ohne Vorauthentifizierung möglich.

Empfohlene Sofortmaßnahmen
  • 1. Perimeter-Inspektion: Relevante Portfreigaben und exponierte Endpunkte unverzüglich scannen.
  • 2. Patch-Applikation: Hersteller-Hotfix einspielen oder betroffene Daemons in isolierte DMZ-Segmente überführen.
  • 3. Telemetrie & EDR-Alerts: Prozessaufrufe und Child-Processes auf anomale Shell-Spawns überwachen.
🔗 Semantisch verwandte Zero-Days MariaDB 11.7 VEC
Ähnliche Beiträge
🔍 Verwandte News

Auch interessante Nachrichten Design and Implementation of a Slurm-Based HPC Cluster

Thematisch verwandte Begriffe: Design, Implementation, SlurmBased, Cluster · 6 Treffer

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Zum Aktualisieren ziehen
ZERO-DAY CVE-2026-96676 | A vulnerability was identified in Fast FAC1900R 20190827_2.0.2. The impa…
Advisory →
TTS Reader • tsecurity.de Voice
tsecurity.de Icon
tsecurity.de App
Offline-Lesen, Eilmeldungen & 0ms Ladezeit

Installiere tsecurity.de direkt auf deinen Home-Bildschirm für das ultimative Vollbild-Magazinerlebnis ohne Browser-Leisten.

Nächster Beitrag
Themen-Radar & Intelligence Matrix
Echtzeit-Taxonomie nach Angriffsvektoren & Plattformen

tsecurity.de Live Threat Radar

🔴 LIVE RADAR
MONITORING
AKTIV
CVE-DATENBANK
LIVE
🔍
Community Radar & Live Chat
Sentinel Bot online • Live-Stream
Dein Cluster: Security Explorer
Match:
lädt…
Verbindung zum Community-Stream wird aufgebaut...
Bearbeitungsmodus — Senden überschreibt deine Nachricht
Community-Puls — was gerade passiert
lädt…
Aktivitäten deiner Analysten
lädt…
Neues Thema oder Eilmeldung einreichen

Reiche interessante Links, Zero-Days oder Debatten ein. Die Community entscheidet per Upvote über die Veröffentlichung.

Heiß diskutierte Einreichungen
🔖 Gespeicherte Artikel
📂 Keine gespeicherten Artikel vorhanden.
Zurück Ziehen Vor
Links: vorheriger Artikel Rechts: nächster Artikel unten: schließen
News NIS-2 Frühwarnung Tier-1 Intel TTP ⏱️ 3 Min vor 10 Min
Artikeldaten werden geladen...

Zurück: vorheriger Vor: nächster
↗ Original-Quelle
Social Reaktionen Deine Reaktion zählt
Einstufung & Relevanz-Poll 0 Stimmen
In sozialen Netzwerken teilen 1-Klick