ERLab is an independent research lab working on reasoning, multimodal learning, autonomous agents, and alignment. We publish openly and build in the open.
Six programs, one goal: making machine intelligence more capable, more understandable, and more trustworthy.
How large language models plan, verify, and self-correct. We study chain-of-thought dynamics, search, and methods that make multi-step reasoning reliable.
Vision–language models that ground language in perception. Focus areas: visual reasoning, document understanding, and efficient cross-modal alignment.
LLM agents that plan long horizons, call tools, and recover from failure. We build benchmarks and study memory, exploration, and credit assignment.
Making capable systems behave as intended. Evaluations, interpretability of internal representations, and robustness to distribution shift.
Getting more capability per FLOP. Scaling laws, data curation, distillation, and post-training recipes for small but strong models.
Reproducible research as infrastructure. We release benchmarks, evaluation harnesses, and open weights whenever we can.
A selection of our latest work. The full list lives in our publication archive.
Our new home on the web, with the full publication archive and open-source releases in one place.
Long-horizon tool-use evaluation, now with per-step credit assignment and a public leaderboard.
Verifiers as Teachers was selected for an oral presentation. See you in Seoul.
We're looking for research interns in reasoning and alignment. Applications close October 1.
A small, senior team of researchers and engineers who like hard problems and short feedback loops.
We hire researchers and engineers who want their work to matter — openly published, openly reviewed, and used by the community.