About · Research
Making uncertainty actionable.
I am a dedicated AI researcher passionate about building trustworthy AI systems that can make the future better. I am currently pursuing my PhD (2023–2027, expected) under the joint supervision of Prof. Elliott Ash and Prof. Mrinmaya Sachan at ETH Zurich, and Prof. Markus Leippold at University of Zurich. I divide my time equally between both institutions.
My research uses uncertainty quantification for two purposes: deciding when a model can be trusted and guiding how it should reason at test time.
Know when a model can be trusted.
A trustworthy model should report low confidence whenever it is likely to be wrong. I extract the uncertainty a model already carries and expose it for human supervision. After calibration, these estimates can support selective prediction and determine when a system should answer, abstain, or defer. This ability to flag its own unreliability is what makes oversight possible.
Guide how a model reasons at test time.
Effective test-time scaling needs feedback on whether the current reasoning is reliable, so it can decide whether further inference is worth the cost. Model-internal UQ can provide this self-assessment without repeatedly querying a separate evaluator. Programmatic verifiers provide definitive feedback when criteria are explicit and cheap to check, while LLM judges cover less formal criteria but add cost and latency and remain tied to a chosen rubric. UQ complements these external signals with a model-native test-time signal for deciding when to continue reasoning, explore another path, or stop.
Projects & community service
Open models, climate NLP, and research communities.
Community collaboration
Apertus: Democratizing Open and Compliant LLMs
Community collaboration on democratizing open and compliant LLMs for global language environments, with contributions to trustworthiness post-training.
Technical report →Community collaboration
When AI Benchmarks Plateau
Community collaboration accepted to ICML 2026 on benchmark saturation and how plateauing scores affect evaluation practice.
ICML 2026 →Workshop
ClimateNLP at ACL 2025
Organizing the second ClimateNLP workshop at ACL 2025, Vienna.
Workshop
ClimateNLP at ACL 2024
Organized the first ClimateNLP workshop at ACL 2024, Bangkok.
Publications
Selected research
Quantifying uncertainty in language models and using it to guide reliable decisions and adaptive reasoning.
ReProbe: Efficient Test-Time Scaling of Multi-Step Reasoning by Probing Internal States of Large Language Models
Reads the model's own internal states mid-reasoning to tell when an answer has settled, scaling test-time reasoning only as far as the model actually needs.
Trace2Skill: Distill Trajectory-Local Lessons into Transferable Agent Skills
Developed at Qwen, Trace2Skill distills successful and failed trajectories into reusable agent skills, using verifier-grounded experience to improve future test-time behavior.
Can Reasoning Help Large Language Models Capture Human Annotator Disagreement?
Tests whether reasoning lets a model recognize when a question is genuinely contested, instead of collapsing real human disagreement into one overconfident answer.
DIRAS: Efficient LLM Annotation of Document Relevance for Retrieval Augmented Generation
Trains small, cheap models to judge whether retrieved documents are truly relevant, giving RAG an external check on its evidence before the generator ever sees it.
Co-DETECT: Collaborative Discovery of Edge Cases in Text Classification
Pairs human annotators with an LLM to surface the edge cases a classifier quietly gets wrong, pointing oversight straight at the model's blind spots.
Towards Faithful and Robust LLM Specialists for Evidence-Based Question-Answering
Trains open QA specialists to answer strictly from the evidence they are given and resist being misled when it is noisy or missing — keeping every answer traceable to its source.
Research Mentorship & Supervision
Skill-RM: Unifying Heterogeneous Evaluation Criteria via Agent Skill
Research mentorship at Qwen · Tao Chen
GD2PO: Mitigating Multi-Reward Conflicts via Group-Dynamic reward-Decoupled Policy Optimization
Research mentorship at Qwen · Haotian Liu
Understanding Failures in LLM Reasoning by Learning Structured Representations of Chain-of-Thought
Research mentee · Tommaso Felice Banfi
Unlocking LLM Legal Reasoning with IRAC-Constrained Chain-of-Thought
Master student · Adam Rahmoun
Full publication list
Education
Academic path
PhD @ ETH D-GESS
ETH Zürich & University of Zürich
MSc in Data Science & Machine Learning
University College London
BEng in Computer Science
University of Hong Kong