Hugging Face releases Carbon-A, a model it used to flag 566 million candidate genes
Hugging Face's biology team released Carbon-A and the Carbon Annotation Database, which it says contains 566 million new gene candidates across 22,617 species. Some candidates were validated in the lab with partners.
Hugging FaceOct 8, 20261 source
LightOn releases LightOnOCR-3, open OCR models that also map document layout
LightOn published LightOnOCR-3 in 0.8B, 1B and 4B sizes. Besides transcribing text, the models output labeled bounding boxes for document regions, describe images and pull numbers from charts.
Hugging FaceOct 8, 20261 source
Hugging Face shows its ML Intern agent training small custom models for a few dollars
In a Hugging Face blog post, staff describe using the ML Intern agent in HuggingChat to plan, train, evaluate and publish small models. One 0.8B model cost about $16 in compute.
Hugging FaceOct 8, 20261 source
TII launches Falcon ASR, a speech model tuned for Arabic and the Emirati dialect
Abu Dhabi's Technology Innovation Institute released Falcon-ASR, a 1.6B-parameter speech recognition model. It reports a 20.92% average word error rate across six Arabic test sets.
Hugging FaceOct 7, 20261 source
NVIDIA says fine-tuned Nemotron models reached gold-medal level at IOI and IMO 2026
NVIDIA reports that specialized versions of Nemotron 3 scored above the gold thresholds at both the 2026 International Olympiad in Informatics and the International Mathematical Olympiad. The runs were unofficial.
Hugging FaceOct 7, 20261 source
llama.cpp adds support for decision models that score options instead of writing text
The llama.cpp server now supports decision models through a new endpoint. These models read an input once and return a probability for each option you give them.
Hugging FaceOct 2, 20261 source
Ai2 open-sources AstaBrief, a small model for writing cited research reports
The Allen Institute for AI released AstaBrief 8B, which turns a research question and retrieved literature excerpts into a cited report. It powers the Fast mode of report generation in Ai2's Asta platform.
Hugging FaceOct 2, 20261 source
Hugging Face launches an open, multilingual text-to-speech leaderboard
Hugging Face and collaborators launched the Open TTS Leaderboard to evaluate open-source text-to-speech and voice-cloning models across languages. The authors say open models are underrepresented on arena-style rankings.
Hugging FaceSep 30, 20261 source
H Company releases Holo4 models for agents that operate software
Paris-based H Company introduced Holo4, agentic models in 27B dense and 35B-A3B mixture-of-experts sizes, plus an updated Holotron4 Nano. The models work through GUIs, code, MCP and APIs.
Hugging FaceSep 28, 20261 source
Hugging Face Hub adds a home for reinforcement learning environments
Hugging Face added an RL Environments filter to the Hub so agent training and evaluation environments can be shared as dataset repos. It works with frameworks including Harbor, Verifiers, OpenEnv and NVIDIA NeMo Gym.
Hugging FaceSep 28, 20261 source
ESPnet releases YODAS v3, a 1.1-million-hour open speech dataset
The ESPnet team published YODAS v3, an open dataset of about 1.1 million hours of real-world speech in more than 100 languages, with 48kHz audio and timestamped transcripts.
Hugging FaceSep 27, 20261 source