Natural-language guidance · Uncertainty ETH Zürich × University of Zürich

More capable AI. Better human oversight.

I study how to make LLMs and agents trustworthy when human instructions are incomplete and success is hard to verify.

A language model's reasoning runs as a steady signal until it hits a spike of uncertainty, where a human reviewer steps in to check it.

About · Research

Research vision.

I am a PhD candidate under the joint supervision of Prof. Elliott Ash and Prof. Mrinmaya Sachan at ETH Zurich, and Prof. Markus Leippold at University of Zurich. I divide my time equally between both institutions.

I study how to make LLMs and agents trustworthy when human instructions are incomplete and success is hard to verify. My research evolves natural-language guidance for more reliable decisions and uses uncertainty to understand when model judgments deserve trust. My long-term goal is to make advances in AI capability translate into better human oversight. I want models to help us recognize where our requirements need clarification and when they lack the evidence to act.

Selected work

Research highlights.

Natural-language guidance · Judgment

Co-DETECT: Refining natural-language guidance for judgment

Uses annotation uncertainty to surface edge cases, then induces generalizable rules for human review, refining incomplete natural-language guidance for more reliable judgments.

Natural-language guidance · Action

Trace2Skill: Learning natural-language guidance for action

Uses induction across successful and failed executions to improve natural-language guidance for agents, producing reusable skills that transfer across tasks and models without parameter updates.

Guidance · Uncertainty calibration

DIRAS: Calibrating judgments under natural-language guidance

Improves uncertainty calibration through better natural-language guidance, then distills calibrated relevance judgments into smaller models for more efficient deployment.

Uncertainty · Reasoning

ReProbe: Using internal uncertainty signals to guide reasoning

Reads internal uncertainty signals from a frozen LLM to estimate reasoning-step credibility, enabling efficient step verification and test-time search without a separate large process reward model.

Projects · Talks · Service

Community collaboration and services.

Collaboration

Apertus: Democratizing Open and Compliant LLMs

Community collaboration on democratizing open and compliant LLMs for global language environments, with contributions to trustworthiness post-training.

Technical report →

Collaboration

When AI Benchmarks Plateau

Community collaboration accepted to ICML 2026 on benchmark saturation and how plateauing scores affect evaluation practice.

ICML 2026 →

Invited Talk

Keynote at the AI4Law Workshop, ICML 2026

Keynote speech at the AI4Law workshop at ICML 2026 on uncertainty-aware LLM reasoning for legal tasks.

Slides →

Publications

Selected research

Natural-language guidance and uncertainty for more reliable judgments, agent actions, and human oversight.

Education

Academic path

Mar 2023 – Present

PhD @ ETH D-GESS

ETH Zürich & University of Zürich

Sept 2021 – Sept 2022

MSc in Data Science & Machine Learning

University College London

Sep 2017 – Sep 2021

BEng in Computer Science

University of Hong Kong