Author: DigitalOcean - Bewertung: 1x - Views:12
In this video, Yash Sharma, Developer Advocate at DigitalOcean, shows how you can deploy and serve large language model (LLM) using multiple GPUs in managed DigitalOcean Kubernetes cluster which enable efficient and scalable inference for production-ready AI workloads.
Checkout the LLM deployment steps on GithHub -
https://github.com/do-community/DO-Labs/tree/main/nim-deploy/cloud-service-providers/DigitalOcean/DOKS
00:00 - Intro
00:21 - Live Demo
00:59 - Creating cluster with GPU nodes
01:35 - Architecture Diagram
03:22 - Monitoring
03:44 - Closing
To know more, join us on Discord: https://discord.gg/q86bZUbA
To learn more about DigitalOcean, https://www.digitalocean.com/
Follow us on Twitch: / digitaloceantv
Follow us on Twitter: / digitalocean
Like us on Facebook: / digitaloceancloudhosting
Follow us on Instagram: / thedigitalocean
We're hiring: http://grnh.se/aicoph1