Pretraining data acquisition and quality
Since May 2026, I have focused on acquiring and assessing data for pretraining.
- Built acquisition from scratch to an average of multiple billions of raw HTML pages per day, alongside PDFs and other information-rich documents.
- Moved from on-demand fetching to offline batch processing; improved crawl success from 65% to over 90%, including through page rendering.
- Combined domain quality scores with lightweight models for URL scoring.
Agent harness and multi-agent systems
From May 2025 to May 2026, I built the deep-research agent harness from scratch and led its core architecture as development grew from a solo effort to four or five contributors.
I designed multi-source verification to check research claims against external evidence, and enabled parallel tool use and multi-agent coordination. I also contributed to MiroFlow and MiroThinker through framework development, evaluation, and open-source preparation.
Current Public Work
- Apodex 1.1 — Research agents that work with files, data, code, and tools. Public model: Apodex-1.1-mini. FrontierAgent harness.
- Apodex 1.0 — Verification-focused agents for deep research. Public model: Apodex-1.0-mini.
Earlier Projects
- MiroFlow — Agent framework. github.com/MiroMindAI/MiroFlow
- MiroThinker — Agentic model. github.com/MiroMindAI/MiroThinker
- MiroMind-M1 — Foundation reasoning model. github.com/MiroMindAI/MiroMind-M1
Earlier Publications
- MiroThinker: Pushing the Performance Boundaries of Open-Source Research Agents via Model, Context, and Interactive Scaling — Technical Report 2025
- MiroFlow: Towards High-Performance and Robust Open-Source Agent Framework for General Deep Research Tasks — Technical Report 2026
- MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization — Technical Report 2025
- 100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models — Technical Review 2025