AI researcher & engineer · BAAI

Making foundation models work in specialized domains.

I build open data, post-training methods, and agentic/multimodal systems that connect research advances with demanding real-world knowledge and reasoning tasks.

让基础模型真正进入专业领域:开放数据、后训练方法与 Agent / 多模态系统。

Portrait of Xiaofeng Shi
ICML 2026 MechVQA / MechVL
AAAI 2025 CareBot
COLING 2025 MoSLD
3 BAAI collections Led IndustryCorpus, IndustryCorpus2, and Industry Instruction

Signature research

Three connected research directions

My work spans the full path from data and post-training to systems that retrieve, reason, and operate in specialized domains.

Industrial multimodal reasoning

MechVQA / MechVL

An ICML 2026 benchmark with 3.3K mechanical drawings and 21K question-answer pairs, plus a domain-specialized multimodal model trained with SFT and self-play RL.

Agentic knowledge systems

SPAR & SciSage

Multi-agent systems for scholarly retrieval and scientific survey generation, paired with SPARBench and SurveyScope for systematic evaluation.

Domain LLM infrastructure

Open data & post-training

Multilingual industry corpora, instruction data, domain models, data-quality models, and methods for capability-preserving adaptation.

Open resources

Industry data and models

I led the development of these BAAI collections and work to release reusable datasets, models, evaluation assets, and training recipes.

IndustryCorpus2

Multilingual, multi-industry pre-training data with DataRater and classification models.

IndustryCorpus

Open pre-training corpora spanning finance, medicine, law, education, and technology.

Selected publications

Recent work

Complete and current bibliographic records are also available on Google Scholar, ORCID, and OpenReview.

2026
Wnuan: Staged Post-Training for Question Answering over Proprietary Enterprise KnowledgeFirst author / equal contribution
ICML 2026
MechVQA: Benchmarking and Enhancing Multimodal LLMs on Comprehensive Mechanical Drawing UnderstandingCo-first & corresponding author
2026
RAFT: Data Refinement and Adaptive Distillation for Domain Fine-Tuning with Alleviated ForgettingCo-first & corresponding author
2026
ChartWalker: Benchmarking the Cross-Chart RAG Task with Hierarchical Knowledge GraphsCorresponding author
ICML 2026 RLxF
Closing the Feedback Loop: From Experience Extraction to Insight Governance in Verbal Reinforcement LearningCo-author
2025
Rethinking Supervised Fine-Tuning: Emphasizing Key Answer Tokens for Improved LLM AccuracyFirst author · formerly SFTKey
2025
SPAR: Scholar Paper Retrieval with LLM-based Agents for Enhanced Academic SearchFirst author
2025
SciSage: A Multi-Agent Framework for High-Quality Scientific Survey GenerationFirst author
AAAI 2025
CareBot: A Pioneering Full-Process Open-Source Medical Language ModelOpen model, data, and training process
COLING 2025
MoSLD: An Extremely Parameter-Efficient Mixture-of-Shared LoRAs for Multi-Task LearningParameter-efficient multi-task adaptation
2024
CCI3.0-HQ: A Large-Scale Chinese Dataset of High Quality Designed for Pre-Training Large Language Models500GB high-quality Chinese pre-training corpus
2024
Aquila-Med LLM: Pioneering Full-Process Open-Source Medical Language ModelsBilingual medical LLM and training resources

About

Research with an engineering path to adoption

I am an AI researcher and engineer at the Beijing Academy of Artificial Intelligence. Before BAAI, I worked at ByteDance and Meituan. My work has evolved from computer vision and OCR to domain language models, post-training, AI agents, retrieval, and multimodal reasoning.

I aim to release the code, models, datasets, evaluation assets, and training recipes behind my research whenever possible.

More about my background · Technical notes archive

Identity & contact