Vasanth Mohan at AIS5 on Premium Inference, AI Agents & Post-GPU Infrastructure
IgniteGTM · 2026-05-19 · 15м 10с · 19 просмотров · YouTube ↗
Топики: durable-execution
Аудио ещё не скачано.
📝 Summary
Summary ещё не сгенерён.
📜 Transcript
Transcript ещё не сделан.
⚙️ Pipeline jobs
Нет job'ов в очереди.
📄 Описание YouTube
Показать
📍 Recorded live at AI INFRA SUMMIT 5, Plug and Play Tech Center, Sunnyvale, California AI infrastructure is entering a new phase where speed, latency, reasoning performance, and token generation efficiency are becoming just as important as raw compute scale. At AI INFRA SUMMIT 5, Vasanth Mohan breaks down why the next generation of AI infrastructure will move beyond GPU-only architectures and toward specialized inference systems optimized for agentic AI workloads. The talk explores how premium inference, reasoning models, coding agents, and long-running autonomous workflows are reshaping data center design, memory architectures, and AI infrastructure economics. Vasanth also explains how disaggregated inference architectures separate prefill and decode workloads across specialized chips, and why future AI infrastructure stacks will increasingly combine CPUs, GPUs, and inference accelerators together. Topics include: • Premium inference and fast-token AI systems • Agentic AI workloads and reasoning-token growth • Why latency is becoming critical for AI infrastructure • Coding agents and long-running autonomous workflows • Disaggregated inference architectures • Prefill versus decode optimization • Memory architectures for trillion-parameter models • Energy efficiency and throughput optimization for inference infrastructure 📣 Super early bird tickets are now available for the next AI INFRA SUMMIT → https://luma.com/aiinfra6