Author: Microsoft Developer - Bewertung: 0x - Views:11
Fine tuning, serving, and inference are core components of real world AI adoption. In this panel Rob Ferguson, Daniel Han (Unsloth), Mark Saroufim (Stealth Startup) and other practitioners discuss how teams are customizing models, taking them to production, and running them efficiently at scale. The conversation covers tradeoffs in fine tuning techniques, infrastructure choices for serving, optimizing inference cost and latency, and what’s working today for builders shipping AI powered products.
Seating for this session is first-come, first-served. Add it to your schedule to plan your day and arrive early to secure a spot.
To learn more, please check out these resources:
* https://aka.ms/build26-next-steps
𝗦𝗽𝗲𝗮𝗸𝗲𝗿𝘀:
* Rob Ferguson
* Daniel Han
* Mark Saroufin
𝗦𝗲𝘀𝘀𝗶𝗼𝗻 𝗜𝗻𝗳𝗼𝗿𝗺𝗮𝘁𝗶𝗼𝗻:
This is one of many sessions from the Microsoft Build 2026 event. View even more sessions on-demand and learn about Microsoft Build at https://build.microsoft.com
BRK234 | English (US) | Working with models
Breakout | (200) Intermediate
#MSBuild
Chapters:
0:00 - Overview of DeepSeek R1 and reinforcement learning infrastructure challenges
00:13:02 - Historical context of reinforcement learning applied to optimization tasks like video compression
00:14:02 - Discussing major RL challenges — defining correct reward functions for diverse systems
00:19:17 - Debate on scaling reinforcement learning environments toward AGI and limits of the approach
00:25:20 - Efficiency and collaboration example: LoRA with Thinking Machines
00:27:02 - Reinforcement learning perspective and model knowledge debate
00:30:22 - Deep dive into GPU optimization, fusion, CUDA graphs, and kernel writing
00:37:36 - Reflection on human creativity and AI discovering optimization techniques like gradient checkpointing
00:43:05 - Mathematical considerations for softmax limitations and compaction in long contexts
SOCIAL SHARE CARD GENERATOR