Apriel-Reasoner: RL Post-Training for General-Purpose and Efficient Reasoning
Reproducible multi-domain RLVR recipe for a 15B open-weight model, with adaptive domain sampling and a difficulty-aware length penalty for stronger, shorter reasoning.
Research Scientist · London, UK
I'm a Staff Research Scientist at ServiceNow AI Research, based in London. My research centres on RL, LLMs, and scalable distributed systems. I led Apriel-Reasoner and MosaicLeaks, co-authored PipelineRL and TapeAgents, and was a core contributor to the Apriel model series. Before turning to LLMs, my research focused on core reinforcement learning, including implicit offline RL and functional regularization of target networks.
Before ServiceNow I was a Senior Research Engineer at Element AI. I care about scalable and verifiable research: reproducible training recipes, evaluations, and open-source systems other people can build on.
Updates
Selected publications
Reproducible multi-domain RLVR recipe for a 15B open-weight model, with adaptive domain sampling and a difficulty-aware length penalty for stronger, shorter reasoning.
Asynchronous RL infrastructure with in-flight weight updates for fast long-sequence generation while keeping training data near on-policy.
Benchmark and RL framework for agents that must balance task success with privacy leakage from external research queries over multi-hop local and web evidence.
A 15B supernet that supports multiple decoding speed-quality presets from one checkpoint, with released models, serving code, and placement tooling.
arXivTape-centered agent design for resumable state, debugging, evaluation, fine-tuning, prompt tuning, and reusable agent traces.
arXivEarlier work spans implicit offline RL, target-network regularization, active learning, and practical ML workflows for high-stakes investigation settings.
Research systems
Distributed asynchronous reinforcement learning framework for long-horizon LLM training, with in-flight weight updates, multi-domain rollouts, tool use, and scalable post-training workflows.
Framework for building, debugging, serving, and optimizing LLM agents through structured, replayable tapes that connect engineering traces back to model improvement.
Evaluation and training environments for long-horizon LLM agents, providing reproducible task suites and interactive environments for measuring and improving agent behavior.
Blog posts
Mentorship
MosaicLeaks: Privacy Risks in Querying-in-the-Open for Deep Research Agents
Multi-domain agentic RL generalization with CUBE and PipelineRL In progress
Scout Before You Route: Attempt-Conditioned Expert Routing for Software Engineering Agents In progress
Memory-based agents Early research
Research
Engineering
My background combines applied AI research with production software, distributed systems, networking, and infrastructure. I am most interested in research ideas that can be made concrete: implemented, measured, debugged, scaled, and released in a form other people can build on.