Qwen3.8-27B & How to Serve it Fast
Sam Witteveen · 2026-08-18 · 18м 0с · 52 140 просмотров · YouTube ↗
Топики: __inbox__
Аудио ещё не скачано.
📝 Summary
model=deepseek-v4-flash · prompt=summary-v7 · 5 670→15 359 tokens · 2026-08-20 05:02:23
🎯 Главная суть
Через неделю после анонса гигантской Qwen3-Max (2 трлн параметров, доступна только через API) Qwen выложила открытые веса Qwen3-27B. По бенчмаркам модель почти не уступает моделям в разы крупнее и при этом реально запускается локально на обычной GPU. Главный практический вывод: решает не только выбор квантования, но и настройка serving — спекулятивное декодирование и правильный Docker-образ поднимают скорость генерации с ~40 до 100–200+ токенов/сек.
Что даёт Qwen3-27B относительно предшественников
Qwen3-Max с 2 трлн параметров мало кто сможет запустить локально, поэтому выхода открытых весов 27B ждали. По бенчмаркам самой Qwen это существенный скачок почти во всех категориях относительно Qwen3-8B — модели, которая была «народным любимцем» для локального кодинга и локальных агентов. Для тех, кто уже использует 8B, 27B выглядит как прямая замена без потерь. Особенно силён прирост в vision-задачах: computer use и browser use, где 27B обходит Opus 4.6 Max, хотя у Opus именно эти сценарии в фокусе. Независимый Intelligence Index от Artificial Analysis для 27B даёт «безумный» результат: модель почти не уступает GLM-4.5 и сильно опережает Qwen3-32B. По агентному индексу 27B обходит GLM-4 и часть моделей GPT-4.6 — при том что это открытая модель с нормальной скоростью на локальном железе. Сама GLM-4 остаётся хорошей: она обучена скорее на обобщение, чем на погоню за бенчмарками, поэтому её стоит проверять на своих реальных задачах.
Доступные версии: квантования, финтюны, abliteration
На Hugging Face в семействе Qwen3 лежат Max (2T, только API), 8B и 27B в разных квантованиях: полная BF16, FP8, FP4, NVFP. Поверх этого сообщество уже выпустило множество финтюнов — обучение на разных датасетах даёт новые результаты; есть abliterated-версии, первую сделала команда Black Forest AI, следом появилось ещё несколько. Для Mac без NVIDIA GPU MLX-сообщество тоже выпустило свои версии. Выбор конкретной версии и квантования — один из ключевых факторов того, будет ли модель хороша именно для вашего сценария.
Уровень thinking определяет расход токенов
Тест: генерация сайта wellness-ретрита. Без thinking модель справляется вполне прилично. На low получается другой сайт, тоже хороший, с ~500 thinking-токенов. На medium результат похож на low, но парадоксально токенов уходит меньше — тест прогоняли 4–5 раз, стабильно. На high модель потратила ~20 000 thinking-токенов и часто упиралась в лимит, не доводя страницу до конца; содержимое рассуждений при этом во многом повторяет то же, что и на других уровнях — модель «влюбляется в собственные мысли». Поэтому на high нужно закладывать большой контекст и время ожидания, но не гарантию лучшего ответа.
Оверфиттинг на популярных примерах: красный дракон
На тесте SFC-приложения из набора Simon Willison расход токенов различается радикально: low — 420 токенов, medium — ~2000, high — 45 000 (из них 21 000 на рассуждения). При замене объекта с пеликана на красного дракона модель без thinking выходит «детской» и сползает к велосипеду из известного примера; на low — чуть лучше; на medium — хороший результат на велосипеде, но регресс на драконе; на high — результат лучше, но не радикально. Больше thinking даёт прирост, но ценой кратного расхода токенов; medium выглядит оптимальной точкой баланса.
Скорость serving: от базовых 40 до 100+ токенов/сек
Исходная BF16-версия без спекулятивного декодирования давала порядка 40 токенов/сек. Подключение спекулятивного декодирования через vLLM с встроенным MTP-модулем (multi-token prediction) подняло скорость до 100+ токенов/сек. Abliterated-версии в тестах застревали в повторяющихся циклах рассуждений, так что с ними стоит быть осторожнее. Unsloth-квантизация показала около 100+ токенов/сек — хороший вариант, но в финальной связке работала хуже.
Лучшая конфигурация: Docker + NVFP + спекулятивное декодирование
Максимальную скорость дала сборка из официального cookbook Qwen: Docker-контейнер, NVFP-квантизация, спекулятивное декодирование, 128K контекст. На прогоне 6000 токенов средняя скорость составила 173 токена/сек, пики — 200+ токенов/сек; на длинном 128K контексте — около 70 токенов/сек. В ходе экспериментов перебраны vLLM и SGLang, итоговая конфигурация — Docker-образ из cookbook. Совет: под свой сценарий стоит прогнать через кодинг-агента разные версии и уровни thinking — конфигурация влияет на результат сильнее, чем выбор самой модели.
Текущий вердикт и что дальше
На данный момент, если есть возможность засервить эту модель, Qwen3-27B — лучший кандидат для локального AI: это подтверждают и официальные тесты, и собственные эксперименты. Впереди — ожидания от «thinking-cap» финтюнов: версий, которые дают сопоставимый интеллект при меньшем числе токенов на рассуждения, DPO-финтюнов и ещё более быстрого инференса.
📜 Transcript
ru · 2 440 слов · 58 сегментов · flagged: word_run (2 dropped, q=0.95)
Показать текст транскрипта
Okay, so last wey Qen drop the Waites for Qen-3 Max. and It's Undernable. This Modul Is Predy amazing at walking do. Bat the kate Problem with his is it's Two Point for Trillion Parametres. Sether very fue people thergena beaber run this localy ativerms waiting for with the modul the dropt on Friday. This is the twen treepoint EtentyB. So lis vido yena Look at the modul and the Bench Max Centra By Mo Importly I Wanna talk About with Version of this modul. You Shub be running and how, you Shud be running it. Now I Talk Love on the Channel about thwentreepoint 627Bat rially in miny ways that was o the darling modul of a lotov people, douing local coding, running local Agence Centra. A weviving look tet thinks like thinking cap with with a fine tuno of the treepoint six modal. — To besidly get et to be mos think in its rising talkins.usamiv intelligents.ivan Lastwit We looka Glima Thirtyb If You Look Atha Bench Max the acchuly came Out frothis From Queen. We and set this Modulis Doing Substenci bet a the Mato's Music Lima Model in hear.pointsix Model. Now I mention the time, the metow wis publibly Russian Out Noing that the cuin treet ewenty Sevenb was Ganta Drop any day now. A wangena go frue the Benchmax and Hava look at this Most to this video, I wanna look with version of the modul. running Bis o ready the a lots of Different Findunes of the modul. The olsol latsed Different countications of the modul. and Paps more inportly, how shot you be Sating the modul up.runs Localy on urystem. I sofi comining Look A Quen's Own Bench Max. You King see that the comments I made back thing about the Muse Glimah Wha predymuch True this Mollomush sapas Mus Glimat. Muse Glimmer I ting It still every good modol. If in train, a lot more, for generalization then the Actul Bench Max — To still With testing for your Actul use cases. En like a seating tha video I feeld at that Team Is Just ciding stardes. Going forading the future by back to twent 3. So if You look at this, It's a Substential Bump, over Treepoint. & Preemuch Every Erea Heart. Sophi urapy using treepoint. Your Differenty Igenagena Nice troping replacement hear. Ontoma that thesinguperothefa into the Vision Performence in Hera — Amigha Evanthings Like Computerus and stopit that. This is actuly beading Out opus ForpointSixMax. Wish to be fae the sprivis opus Molls When Havally Focusis Much On Computerus and Browser Yuse as the New versions of feel. rete serjustasabin Recording ofofficial Nailis Has relist the Intelligence Index Scores for this 237b modul. In a gota say the Resolt of Prety Insaine. This Is scoring with put not foba hine moduls like tha Gillym Fiveto,5. Ensolgues way a head of thwent 37B Model.Much Any of Open Modul bet U Even Have a Chance of running with ava a Serice Investment in the Hadway. The Athing Is wen You Look at the agentic Index Here. Thisisactuly beading Out. GLMintoo Beating Out Summer of the GPTointSixModules. Again, this is Pretyan Sains For modul the You Can round localy, and geta Deesend Talking Speed Atuba. Tha the keafing I tink is relly Intusting here is With Version of the Modul Duaction Pick. Noth You coming to the wentripountet Family on Hugn Face. Wo kin see the Toopoint for Trillion Parami Modolstap.blow sixteen Version, sels Conace Fool Resolution version and fgurting FPA Version On Thop of thather Oranifoby Versions Out from and Slove. Using NVFP Quantization ansh rembat nologipuse Connector Runlis Quantalization Is well. An Antopov this Therady Lots of Different Findoons of this Modul Oredy Autia Away Peover by traing on Differentadisets to Get New Resoltadavet. Пепл umakings of Unsence Digins of the moduls.people of Dunnet.ru Finduning and the ofer eexamples of wapel of Dunnet.ru Abliteration. So the black Frost AI Team Is won at the First Groups to do this Batinseners Multipol Abliterated Vigions Outher. Aneforseve You don't Have Axes Tu and AMDORNVIAGPU and ya ranning Ona Mac, The MLX Cmmunity Hazordy Relice Multibur versions oftry2 For MLX.ONS, Hones for betons wolls MLX Versions of the Different Crines of Containsation People of Triders well. Thorning doos got of fue the Different versions, harbine Testing of the Weekend more importly agena show you, that axuly how you Sed Up the reasing Is gona be wanna the key Things the Detomens with a this is a good modul, orabad modul for you. Okay, so I fing traning ut Five Different versions of this Modul.raning am ondel Ttoo Pro Max. With Is cynly Sponsor the Compute Hor this, So Olove this Ractune Running on and RTX Pro 6GPromin Video and Hear the apprety bFu. Sokeintex Gitt av ram Hear to play with. So mypatic kayse I Have no poblems bing abil to lod the modul. Rid. Luxury minipuple, Notkanabiaber to lod the foolseenbit. B float Resolution Version of this Modol. FP Version In hea.ria and Sloth Quantisation. Moduls a y Testard Is wow. Suforstup what ever modul the Yupic? You gona realise, thit It's not chust about the conthisation The Totemens with a this Modoles good on. Salamishow utle simple exassise. So heris my standed shtml Test. Were I askay to basicly make, a webside. The Jokers Tario Wellness retreat. From Byoce Heart that this is a version of the webside with made with no thinking at all in here. So, it's Donna Prely good job. Just basicly thinking tand off, complitly. You cansirue me coming me look a tha resoutia the a no thinking Talkings raditisco streat in too generation for this one. If You look at the samething for reasing on with a low thinking. You concewet cryd a different webside. Geting something the Stulvery Nice this olotov effecting It. and you mayace if and like the version bat, a now thinking and alling Hear. Get It is the the Different levels of thinking Consuing Susi Amounts Different talking. So her, wook at tha co, wee cand se we got fivehundredwell thinking talkins. An a seams to be reasonly consistent win I wrannel a furedifrentimesly get cuatabit less, and cazzenly get's more.pans of tos deracti asing et to doo, bat here, low is giting five hundred ent welf. Sophi coming here and look the Midium Thinking You Can see that this is china Symular axuly to the Low One right. Say wood presume hera the thisis gun acculy be using more thinking talkins. Bat axuly it ins up using less. I frundisor five Times I not show why onlis Paticuly Task Medium Thinking Seems to be using Less. Then Love Thinking here. Now. The Question Shood be well what About high well. The ally Hab hy, the have xigh thinking. And this is with the modulgos nots. You see heat that I don't have a finish webside to show you on this One. Bis the Isha hero is the wall I liminter to the to key Max Talkings Out. I achaly runout. Bis I spen 7talkins On Thinking. Now. I wrunnis multibul times, trind to get one a axuly finish. And I have thinking Talkins, bashy aswenty 20ouson Thinking Talkins, anif y look atham. Icosin Twins Sain Detailingher So walt auchuly producers, a lot of the Simo Serical Stuff as the ether Ones — Itust Talkswitself I Love About this. This is a General Consistenting by yove Notist Is That If You gota Runes Model onxhi, You gota be pepa the One you wanta make the contexwindow. Werry big Butoo olsy Gand a be seling Around waiting for the Thinking Talkings, too actuly proces.runing on the fPates with Dundisixamples. But Ivalse tride his on the 6 Modul and on the unsloth Modul. For Eate of this on X Hy Thinking You get in Sainy Long amounts of Thinking. Now, if You Look At tha Sainthing for the SFC Tess. So this is the Generate a Palaken SFC Aptekanis From Simon Wilsens Tess Anikasy this Is on X Hies witha a leventhousand Talkings of Thinking To Get this. Simmula One of Defferenty staling to Miss Som a the Deatales dersh of if We tarn rising off, and just hava look the Street One Out with no thinking wital Pre Ungly Paliken Oldo One Stage his prove wado Bang consine really good. Me, seams to be having the rising on Medium. A just show you cookly hammat it over fiting on the Pulican example, here I change at to a red dragon. And you can see that it Stalking at the bike ry on no talkings out bat a red Dragon is cana Childish. If We Go Up to love thinking. Something get the bit bet. If We go up to medium Thinking, Ukreeb insiderets crynof regress obit on the Dragons — Stopery Good on the bysico, by non on the Dragen. Ambodyway thistime the low Thinking Actory is cry Loads, 4hundreen Twenty Talkins As apost to Medium.tohthouson Talkings — Heart for the thinking. Nakes me go to Xy wh husing ferty fivetouson Talkins Awish 2 with a thinking Is a betta. Ye, it's prow betta, he and think I with say it is grete. City Backscape, with to Pubigut a peta bysicle. So Yes. will get you better resolts ye man talking this is cansain. Say concie Everlo Hattowsen the Talkingwindow and It Used up trety Fivivetouson Talkinswish Twenty1touson Hear — Wat Thinking Talkings Just to Get The Dragonwitch. Gesis okay, looking bot assely do great looking hear. Brights Re never ino seling the Number of Thinking Talkins, Is raillian Port. Vea thing is with version of the modal Duaction go for So first of I Tested the basic b floatsixteen Original Version. And that wis gilling me sumwer round for twty talkingspurs second. Purwar generation. Now. A that pwain I hat no speculative decoading going on. From the be kamePrety Cleer The U Wanto basicly Have the Speculative Decoading on epovivan Ex Trees. Multitalking Production for Tree Talkings Tha Nell Longot Me A Good Speed Bump in acti ranning the model. Never running Mosly Ibin Running on VLMStoo WozFersion of the modal From Queens Anispeckudikoding with Geaning me run edione handredy Talkins pusecon. Competue the be floot 6 Version Agentis with using the bill in MTP Model, for peculative De coading. and ef you goot at set up ride villerm candust Take Kare that for you. Now Around Is Time I Olso tride Out Summo the abliterated Versions of the modul. Am wan Fund am intusting and yes, tay sel of opentings Up. Repited Loops on the Thinking and just kit Stuck in that Repited Lup. So onese prebiat a moment, Own bothy with using the abliturated Ons. Nextop wath — the Uns Love modul. wis head rilly good talking reats with thp as sgeling Around Hundrenteny Talkings fu Second hear Is very Nice countaisation of the modul, Afan Warkrol Lotov the Differents.Rasc Langs Version of Sevings Modul Sothe not Saporting Old the Different Hard way, the realy Focus a lot More on the black Well GPS. Configuad for Doing Art Expressixer FPND Sput and a Tuns Out that Takne Everythings of the JP and loading tat in a Dockar Configuration was Able to Get Me very for speeds. I flook the Xaccount. Theakana Showing wits On the Rickout Ride The Clemi Head inged Up to 20o hundreds Talkins Pur Second And the dou tat the using a wray Paticula NVFP Modul I found the win I use the Sed Up with the Unslove cuantalsation It iding workis wow. a modul waits Of I'm Geading Much Fasta Talkings puseconds. Suckis hea this estet 6touson Talkens. With and everage of hundred seveny 3alkings for second I case for Thinksa not too long Like This. I can of tan Axy getroundoundredoundredy Talkinsecon. Now I with say the everage Is probably a betunderdalkingsecond Browndredmaby Little Bibelothater that Speed that's flyine. I gold running in adoha cantainer, with Is using the Paticular Image from the cookbook. It's the NVFor Cuant from Them, with the Deespak Speculative Decoding India and scotor footounredskontexwindow Running beautifuly India. Upstoging Everage Speedseveny Talkins Pursecond The OveralJ Stefinishup I wissay urilany Too Experiment with Different Versions of the Modul to see Howscana work for epeticula Yous Cases Onesive Gottap modul — The Differentonfigurations of the A Mount of Resing Talkins, that You Wanto Set. Andreventens at what infrence Lide rew You want You. Sow I Tested viom in hea, I Tested SG Lang. Lang with the Wina For this finel configuration in heat if You Got AGP with a Loa Viram, ePubly Osowna Testa Lama CP Alagulsay om morever you doing the Testing Give this Pages of wa people Talking abaut How Ised Ropenstephut bat to your coding Agent and havite Test Out Different configurations to see what Works best for you. For me curently if You can save this modul, this is the modul to be for Loco A. Reauly is doing so well. The Fect as official nalsis intelligens Test, Showses to be so histconess for me. Things I fel using a and testing It Out. and I finshop y saing the the the a tink Itskinget ivan Petter I'm really looking for to say to wiget a thinking caps version of this.usis less Talkings for a samimar intelligents. Other findunes Diraction pruv Intelligence. Thinks Like Fuse Cornols. Be Surfet even Fasta Talkin's Pursecond. Salamenion the Comments — How Your runing this what You can't setup as Harmony Talkings Persecond Iry wisit this Modol. On the Digiects Spacs — of the GB Tens and show you with One too mabie veve and three Machines, whath You can do vere for this Canithing Beclure this ystema Hughestep ford both Notalme in open Waits. Bundy Local AI Way Kan Runny Stangs On Surprosuma Hardwee. If You Like the video, pliceck like on sipscry and I watalkinx video by Funow
⚙️ Pipeline jobs
| Stage | Status | Att. | Updated | Error |
|---|---|---|---|---|
| download | done | 2/3 | 2026-08-20 04:43:16 | |
| transcribe | done | 1/3 | 2026-08-20 05:00:11 | |
| summarize | done | 1/3 | 2026-08-20 05:02:23 | |
| embed | done | 1/3 | 2026-08-20 05:02:25 |
📄 Описание YouTube
Показать
In this video, I look at the long awaited Qwen3.8-27B model. Both what it can do and how to serve it at the maximum tokens per second Thanks to Dell for Sponsoring the Compute #DellProPrecision #DellProMax #DellTech #NVIDIA 📖 Website: https://qwen.ai/ 🤗 HF: https://huggingface.co/collections/Qwen/qwen38 SGLang: https://lmsysorg.mintlify.app/cookbook/autoregressive/Qwen/Qwen3.8-27B Twitter: https://x.com/Sam_Witteveen 🕵️ Interested in building LLM Agents? Fill out the form below Building LLM Agents Form: https://drp.li/dIMes 👨💻Github: https://github.com/samwit/llm-tutorials ⏱️Time Stamps: 00:00 Intro 00:50 ThinkingCap 01:25 Qwen3.8 - 27B 01:59 Different Versions on Hugging Face 02:12 Benchmarks 03:14 Artificial Analysis Benchmark 04:10 Qwen3.8-27B on Hugging Face 06:40 Demo 14:17 SGLang