Special Topics in Kernels, RL, Reward Hacking in Agents — Daniel Han, Unsloth
AI Engineer · 2026-07-17 · 2ч 20м · 3 690 просмотров · YouTube ↗
Топики: ai-agent-orchestration
Аудио ещё не скачано.
📝 Summary
Summary ещё не сгенерён.
📜 Transcript
Transcript ещё не сделан.
⚙️ Pipeline jobs
Нет job'ов в очереди.
📄 Описание YouTube
Показать
An advanced seminar (good prerequisites: Daniel's 2024 and 2025 hit AIE workshops, but all are welcome!) PLS WATCH: https://www.youtube.com/@aiDotEngineer/search?query=daniel%20han Timestamps: 0:00 Introduction to Unsloth and model distribution 2:32 The State of AI: Meter plots and performance trends 20:26 Open Source vs. Closed Source models 38:51 Throughput maxing and accuracy minimizing 1:03:00 Benchmarking and cheating in AI 1:37:49 Kernels and algorithmic improvements 2:04:16 Reinforcement learning primer 2:05:19 Reward hacking and AI agents Viral Quotes & Pull Quotes: "If you make the model 86% smaller, it does not get 86% dumber... it only gets 14% less dumb." (29:52) "Reinforcement learning is terrible, but everything else is even worse." (20:43) "The model becomes not important anymore; it's the harness or the tool that is actually the most important thing." (38:23) "If a model can finish a task that takes a human 16 hours, can a model finish that task?" (2:49)