← все видео

Special Topics in Kernels, RL, Reward Hacking in Agents — Daniel Han, Unsloth

AI Engineer · 2026-07-17 · 2ч 20м · 3 690 просмотров · YouTube ↗

Топики: ai-agent-orchestration

Аудио ещё не скачано.

📝 Summary

Summary ещё не сгенерён.

📜 Transcript

Transcript ещё не сделан.

⚙️ Pipeline jobs

Нет job'ов в очереди.

📄 Описание YouTube

Показать
An advanced seminar (good prerequisites: Daniel's 2024 and 2025 hit AIE workshops, but all are welcome!)

PLS WATCH: https://www.youtube.com/@aiDotEngineer/search?query=daniel%20han

Timestamps:

0:00 Introduction to Unsloth and model distribution
2:32 The State of AI: Meter plots and performance trends
20:26 Open Source vs. Closed Source models
38:51 Throughput maxing and accuracy minimizing
1:03:00 Benchmarking and cheating in AI
1:37:49 Kernels and algorithmic improvements
2:04:16 Reinforcement learning primer
2:05:19 Reward hacking and AI agents

Viral Quotes & Pull Quotes:

"If you make the model 86% smaller, it does not get 86% dumber... it only gets 14% less dumb." (29:52)
"Reinforcement learning is terrible, but everything else is even worse." (20:43)
"The model becomes not important anymore; it's the harness or the tool that is actually the most important thing." (38:23)
"If a model can finish a task that takes a human 16 hours, can a model finish that task?" (2:49)