
Two things landed that I want to keep in the same frame. NVIDIA and Hugging Face have signed an acquisition agreement, which confirms the read I wrote when the talks were still unconfirmed: the portal was the prize. And IFM shipped real small open-weight models, K2 Horizon at 0.9B, 3.7B, and 7B, that you can actually aim at a device. Do not let Flash or Max branding steal that second slot.
Signed, not closed
I wrote about this when The Information first reported a bid. The thesis then was that a well-positioned developer portal that compounds usefulness is a durable, positive thing. That still holds. What changed is the paperwork.
On September 3, Jensen Huang announced that NVIDIA had agreed to acquire Hugging Face for $12,930,300,000. The company’s 8-K describes roughly $11.9 billion to stockholders and up to about $1.0 billion for retention. The expected close is in the first half of 2027. This is a signed agreement. The deal has not closed yet.
Huang’s post is careful about the platform promise. Hugging Face will remain open to the wider ecosystem. Developers choose their models, frameworks, clouds, and inference providers. NVIDIA compute will not be required to build on or deploy through Hugging Face. Multi-silicon support is part of the same commitment language. I would treat those as live-deal commitments, not as post-close facts we can already measure.
The 8-K also flags a real risk for the Hub: restrictions on China-origin open models. The filing names it as a disclosed issue for a platform whose job is to host weights from everywhere. Neutrality of the Hub is still the thing to watch between signing and close.
I do not read ownership as a eulogy for open source. Hugging Face became the place builders actually go: model cards, datasets, Spaces, the first stop for a student and a startup alike. NVIDIA says it wants to scale that. Whether the public square still feels like a public square after the close is the question the community will answer with its uploads. Until then, the encouraging part is the same as before. The portal was valuable enough that someone was willing to put a precise twelve-digit number on it.
K2 Horizon, in the sizes that fit a pocket
Separately, IFM released K2 Horizon under Apache 2.0. The fleet is six models. The ones that belong in a Small AI conversation are the dense 0.9B, 3.7B, and 7B. The 32B, the 36B-A4B MoE, and the 375B-A23B are different machines. Keep them out of the phone-class drawer.
The 0.9B is aimed at watches and glasses class hardware, with 128K context (YaRN stretch listed to 131,072). The 3.7B and 7B are the phone and on-device sizes. That is the shape I want on a syllabus: permissive license, named sizes, a clear claim about where they run.
IFM reports AIME 2026 at 48.5 for the 0.9B against 0.2 for Qwen3.5-0.8B. That is a vendor number, and the setup used high reasoning effort with long output tokens. It is an interesting lab result. Treat it as a phone product only after someone publishes 4-bit device benches you can reproduce. IFM also disclosed reward-hacking issues on larger siblings in the family. Credit them for saying it. Then go measure the small ones yourself.
When a lab ships a “Flash” or “Max” model with a few billion active parameters inside a hundred-billion-plus total, that is efficiency-tier frontier work. That is a different story from a 0.9B or 3.7B you can put next to a camera roll. Ask for total parameters, active parameters, license, and the machine they ran. The K2 Horizon small dense SKUs clear that bar. The giant siblings do not need to crowd them out.
A short note on African-language translation
Tether’s QVAC line also shipped TranslatePsy AfriSLM: supervised fine-tunes of Qwen3.5 at 0.8B, 2B, and 4B for 19 African languages, with Apache 2.0 weights. That is specialized local machine translation, not a new general SLM family. If you teach multilingual edge work, it is worth a look next to the base Qwen3.5 cards. The synthetic data mix carries a CC-BY-NC 4.0 note, so read the cards before you ship.
What I would do this week
Re-read Huang’s acquisition post and the 8-K summary with the close date and Hub commitments in mind. Then download K2 Horizon 0.9B or 3.7B, quantize it, and run the same prompt you use for other sub-10B models on a machine you already own. Write down latency, memory, and where reasoning helps. If you care about African languages, try AfriSLM on a short FLORES-style sample and compare to the base checkpoint.
A signed portal deal and a handful of truly small open weights can live in the same newsletter without cancelling each other out. The Hub still has to stay useful. The small models still have to run where people actually work. Both are encouraging, if we keep the language honest.
Time to dig deeper
Start with NVIDIA’s Hugging Face announcement, my earlier portal piece, and IFM’s K2 Horizon post. Load one of the sub-10B checkpoints. Leave the 32B-and-up SKUs for another afternoon. The future of Small AI gets clearer when we celebrate the models that fit the device, and watch the portal commitments all the way to close.