Uncertainty Quantification · NLP ETH Zürich × University of Zürich

AI systems that know when to be trusted.

I quantify the uncertainty a model already carries and use it to guide human supervision and adaptive test-time reasoning.

A language model's reasoning runs as a steady signal until it hits a spike of uncertainty, where a human reviewer steps in to check it.

About · Research

Making uncertainty actionable.

I am a dedicated AI researcher passionate about building trustworthy AI systems that can make the future better. I am currently pursuing my PhD (2023–2027, expected) under the joint supervision of Prof. Elliott Ash and Prof. Mrinmaya Sachan at ETH Zurich, and Prof. Markus Leippold at University of Zurich. I divide my time equally between both institutions.

My research uses uncertainty quantification for two purposes: deciding when a model can be trusted and guiding how it should reason at test time.

Improving reliability

Know when a model can be trusted.

A trustworthy model should report low confidence whenever it is likely to be wrong. I extract the uncertainty a model already carries and expose it for human supervision. After calibration, these estimates can support selective prediction and determine when a system should answer, abstain, or defer. This ability to flag its own unreliability is what makes oversight possible.

Enhancing capability

Guide how a model reasons at test time.

Effective test-time scaling needs feedback on whether the current reasoning is reliable, so it can decide whether further inference is worth the cost. Model-internal UQ can provide this self-assessment without repeatedly querying a separate evaluator. Programmatic verifiers provide definitive feedback when criteria are explicit and cheap to check, while LLM judges cover less formal criteria but add cost and latency and remain tied to a chosen rubric. UQ complements these external signals with a model-native test-time signal for deciding when to continue reasoning, explore another path, or stop.

Projects & community service

Open models, climate NLP, and research communities.

Community collaboration

Apertus: Democratizing Open and Compliant LLMs

Community collaboration on democratizing open and compliant LLMs for global language environments, with contributions to trustworthiness post-training.

Technical report →

Community collaboration

When AI Benchmarks Plateau

Community collaboration accepted to ICML 2026 on benchmark saturation and how plateauing scores affect evaluation practice.

ICML 2026 →

Workshop

ClimateNLP at ACL 2025

Organizing the second ClimateNLP workshop at ACL 2025, Vienna.

Workshop

ClimateNLP at ACL 2024

Organized the first ClimateNLP workshop at ACL 2024, Bangkok.

Publications

Selected research

Quantifying uncertainty in language models and using it to guide reliable decisions and adaptive reasoning.

Research Mentorship & Supervision

05
2026 AI4Law @ ICML 2026

Unlocking LLM Legal Reasoning with IRAC-Constrained Chain-of-Thought

Master student · Adam Rahmoun

Full publication list

Education

Academic path

Mar 2023 – Present

PhD @ ETH D-GESS

ETH Zürich & University of Zürich

Sept 2021 – Sept 2022

MSc in Data Science & Machine Learning

University College London

Sep 2017 – Sep 2021

BEng in Computer Science

University of Hong Kong