预训练数据获取与质量
自 2026 年 5 月起,我主要负责预训练数据获取与质量评估。
- 从零搭建数据获取系统,将规模扩展到平均每天数十亿个原始 HTML 页面,同时获取 PDF 等信息密集型文档。
- 从按需抓取转向离线批处理,通过页面渲染等优化,将抓取成功率从 65% 提升到 90% 以上。
- 结合域名质量评分与轻量模型,对 URL 进行质量评估。
智能体框架与多智能体系统
2025 年 5 月至 2026 年 5 月,我从零搭建深度研究智能体框架,并在开发从单人扩展到四五位贡献者的过程中负责核心架构。
我设计多源验证机制,用外部证据核查研究结论,并实现并行工具调用与多智能体协作。我也参与了 MiroFlow、MiroThinker 的框架开发、评测与开源准备。
当前公开成果
- Apodex 1.1 — 面向文件、数据、代码与工具协作的研究智能体。公开模型:Apodex-1.1-mini。FrontierAgent 执行框架。
- Apodex 1.0 — 注重验证的深度研究智能体。公开模型:Apodex-1.0-mini。
早期项目
- MiroFlow — 开源智能体框架。github.com/MiroMindAI/MiroFlow
- MiroThinker — 开源智能体模型。github.com/MiroMindAI/MiroThinker
- MiroMind-M1 — 开源基础推理模型。github.com/MiroMindAI/MiroMind-M1
早期论文
- MiroThinker: Pushing the Performance Boundaries of Open-Source Research Agents via Model, Context, and Interactive Scaling — Technical Report 2025
- MiroFlow: Towards High-Performance and Robust Open-Source Agent Framework for General Deep Research Tasks — Technical Report 2026
- MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization — Technical Report 2025
- 100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models — Technical Review 2025