Publications

(2026). Learning from Use: Test-Time Learning in Large Language Models and Agents. Preprint 2026.

SSRN

(2026). Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical Cases. Preprint 2026.

Arxiv

(2026). PACE: A Proxy for Agentic Capability Evaluation. Preprint 2026.

Arxiv

(2026). PaperMentor: A human-centered multi-agent writing tutor for AI research papers in Overleaf. ACL 2026 Demo.

ACL Anthology Code Demo

(2026). OdysSim: Building Foundation Models for Human Behavior Simulation. Preprint 2026.

Arxiv

(2026). Re-Centering Humans in LLM Personalization. Preprint 2026.

Arxiv

(2026). Knowledge Index of Noah's Ark. Preprint 2026.

Arxiv

(2026). LT2: Linear-Time Looped Transformers. Preprint 2026.

Arxiv

(2026). Reinforcing Human Behavior Simulation via Verbal Feedback. Preprint 2026.

Arxiv

(2026). MixSD: Mixed Contextual Self-Distillation for Knowledge Injection. Preprint 2026.

Arxiv

(2026). Transferability for General Reasoning: An Automated Curriculum for Multi-Domain RLVR. Best Paper Award at ICML 2026 RLxF Workshop.

Arxiv

(2026). Self-distillation zero: Self-revision turns binary rewards into dense supervision. Best Paper Award at ICML 2026 RLxF Workshop.

Arxiv

(2026). CLT-Forge: A Scalable Library for Cross-Layer Transcoders and Attribution Graphs. Preprint 2026.

Arxiv

(2026). Mind the sim2real gap in user simulation for agentic tasks. Preprint 2026.

Arxiv

(2026). Making Complex Reasoning Student-Friendly: A Hybrid LLM-to-SLM Distillation Framework. Preprint 2026.

(2025). Stabilizing Reinforcement Learning for Honesty Alignment in Language Models on Deductive Reasoning. AAAI 2026 Bridge LMReasoning Workshop, AAAI 2026 MATH4AI Workshop.

Arxiv

(2025). LLM Microscope: What Model Internals Reveal About Answer Correctness and Context Use. COLM 2025 Interplay Workshop.

Arxiv

(2025). CauSciBench: A Comprehensive Benchmark on End-to-End Causal Inference for Scientific Research. ICML 2026.

(2025). Taming Object Hallucinations with Verified Atomic Confidence Estimation. EACL 2026.

Arxiv

(2025). CORE: Measuring Multi-Agent LLM Interaction Quality under Game-Theoretic Pressures. EACL 2026.

Arxiv

(2025). Synthetic Socratic Debates: Examining Persona Effects on Moral Decision and Persuasion Dynamics. EMNLP 2025 Main.

Arxiv

(2025). Humanizing Machines: Rethinking LLM Anthropomorphism Through a Multi-Level Framework of Design. EMNLP 2025 Main Oral.

Arxiv

(2025). BIG5-CHAT: Shaping LLM Personalities Through Training on Human-Grounded Data. ACL 2025 Main.

Arxiv Code

(2025). Towards Global AI Inclusivity: A Large-Scale Multilingual Terminology Dataset (GIST). ACL 2025 Findings.

Arxiv

(2025). Uncovering and Understanding Social Media Censorship across Countries. ACL 2025 Findings.

(2025). Chumor 2.0: Towards Benchmarking Chinese Humor Understanding. ACL 2025 Findings.

Arxiv

(2025). EmoNews: A Spoken Dialogue System for Expressive News Conversations. SigDial 2025 Demo.

Arxiv

(2024). Language Model Alignment in Multilingual Trolley Problems. ICLR 2025 and Best Paper Award at NeurIPS 2024 Pluralistic Alignment Workshop.

Arxiv

(2024). Implicit Personalization in Language Models: A Systematic Study. EMNLP 2024 Findings.

Arxiv

(2024). Synatra: Turning indirect knowledge into direct demonstrations for digital agents at scale. NeurIPS 2024.

Arxiv

(2024). Inducing Elasticity in Foundation Models: Post-Training Techniques for Adaptable Inference. NeurIPS ENLSP Workshop 2024.

(2024). Chumor 1.0: A Truly Funny and Challenging Chinese Humor Understanding Dataset from Ruo Zhi Ba. Preprint 2024.

Arxiv

(2024). Automatic Generation of Model and Data Cards: A Step Towards Responsible AI. NAACL 2024 Oral.

Arxiv

(2024). Analyzing the Role of Semantic Representations in the Era of Large Language Models. NAACL 2024.

Arxiv

(2023). Can Large Language Models Infer Causation from Correlation?. ICLR 2024.

Arxiv

(2023). Bias Amplification Enhances Minority Group Performance. TMLR 2023.

Arxiv

(2023). Voices of Her: Analyzing Gender Differences in the AI Publication World. ACL 2025 NLP for Positive Impact Workshop.

Arxiv