Mooncake is a service-layer system designed to support LLM execution by separating the PREFILL phase (initial context construction) from the DECODE phase (token generation).
It leverages CPU, SSD, and DRAM resources to efficiently manage the KVCache generated during prompt execution on vLLM, enabling reuse of previously computed data and reducing...
🛡️ VERIFIED CYBER INTELLIGENCE ID: #3124320