About Bo Peng
I am a Ph.D. student in Computer Science at Shanghai Jiao Tong University (SJTU), advised by Prof. Yu Qiao and working with Dr. Chaochao Lu. I am a member of the Wu Wenjun AI Honors PhD Program and expect to graduate in 2028. I received my bachelor's degree from SJTU's Artificial Intelligence Pilot Program in 2023. My research mainly focuses on Large Language Models, AI Agents and Causality.
I am currently a research intern at Shanghai AI Laboratory (Nov 2022–present).
News
- 2026.08 CauSight is accepted to EMNLP 2026, and CauAudit to Findings of EMNLP 2026!
- 2026.06 CauTion, our framework for deciding when to trust LLMs in ensemble causal discovery, is available as a preprint.
- 2026.05 CauScale is accepted to ICML 2026!
- 2026.04 Our benchmark for multi-day coworker agents, ClawMark, is available as a preprint.
- 2026.01 CauScientist, our framework combining LLMs with statistical verification for causal discovery, is available as a preprint.
Publications
2026
ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents
arXiv preprint 2026
A benchmark of 100 coworker tasks across 13 professional scenarios, with evolving service state, raw multimodal evidence, and 1,537 deterministic checks.
The Illusion of Visual Tool-Use: A Causal Audit of Thinking with Images
EMNLP · Findings 2026
A causal audit of whether visual tool outputs influence model decisions, with interventions at the policy, trajectory, and individual-step levels.
CauSight: Learning to Supersense for Visual Causal Discovery
EMNLP 2026
Visual causal discovery from images through VCG-32K, synthesized reasoning trajectories, and reinforcement learning with causal rewards.
CauScale: Neural Causal Discovery at Scale
ICML 2026
An efficient neural architecture for causal discovery, scaling inference to graphs with up to 1,000 nodes through a two-stream design and shared attention.
CauTion: Knowing When to Trust LLMs for Ensemble Causal Discovery
arXiv preprint 2026
Combines statistical algorithm consensus with reliability-calibrated LLM arbitration to resolve disputed causal relations, followed by cycle repair.
CauScientist: Teaching LLMs to Respect Data for Causal Discovery
arXiv preprint 2026
Combines LLM-proposed causal hypotheses with statistical verification and error memory to guide causal structure search.
2025
IP-Dialog: Evaluating Implicit Personalization in Dialogue Systems with Synthetic Data
EMNLP · Findings 2025
A synthetic-data benchmark and training pipeline for inferring implicit user attributes from dialogue, covering 10 tasks and 12 attribute types.
2024
CELLO: Causal Evaluation of Large Vision-Language Models
EMNLP 2024
Evaluates visual causal reasoning through 14,094 questions spanning discovery, association, intervention, and counterfactual reasoning.
Causal Evaluation of Language Models
Technical report 2024
A framework and benchmark for evaluating language models across causal tasks, prompting strategies, metrics, and error types.
ConditionVideo: Training-Free Condition-Guided Video Generation
AAAI 2024
Training-free video generation using pretrained image diffusion models, with condition-guided motion and improved temporal coherence.
Open-Source Tools
- Paper-Researcher [GitHub]
A Python MCP toolkit for literature retrieval and citation-network exploration using Semantic Scholar, with bounded traversal, resumable checkpoints, and traceable discovery.
Education
- 2023 – 2028 (expected), Ph.D. in Computer Science, Shanghai Jiao Tong University.
Advisor: Prof. Yu Qiao. Wu Wenjun AI Honors PhD Program. - 2019 – 2023, Bachelor's degree, Shanghai Jiao Tong University.
Artificial Intelligence Pilot Program. GPA: 3.83 / 4.3.
Internships
- Nov 2022 – Present, Research Intern, Shanghai Artificial Intelligence Laboratory.