Research vision / Jingwei Ni
More capable AI. Better human oversight.
I study how to make LLMs and agents trustworthy when human instructions are incomplete and success is hard to verify.
Research vision
Connecting guidance, uncertainty, and oversight.
I study how to make LLMs and agents trustworthy when human instructions are incomplete and success is hard to verify. My research evolves natural-language guidance for more reliable decisions and uses uncertainty to understand when model judgments deserve trust. My long-term goal is to make advances in AI capability translate into better human oversight. I want models to help us recognize where our requirements need clarification and when they lack the evidence to act.
These questions are connected. Clearer guidance gives us a more precise basis for judging a model’s behavior. Uncertainty can help identify cases where that guidance is incomplete or the available evidence does not support a confident judgment. Understanding which problem we face matters: clarifying a requirement, gathering evidence, and improving a model’s reasoning address different limitations.
Natural-language guidance
Learning guidance from experience.
I study how observations across cases can be distilled into guidance that generalizes. In Co-DETECT, uncertain annotations help surface edge cases, and induction produces rules for human review. In Trace2Skill, lessons from successful and failed executions become reusable guidance for agents.
The distinction between judgment and action is important. Evaluation criteria describe how to assess behavior; procedural guidance helps an agent act effectively. My future research asks when experience reveals a gap in a procedure and when it reveals a requirement that needs clarification.
Uncertainty
Understanding when judgments deserve trust.
DIRAS connects natural-language relevance criteria with calibrated judgments that can be distilled into smaller models. ReProbe reads a frozen model’s internal states to estimate reasoning-step credibility, supporting efficient verification and test-time search.
These methods provide evidence about reliability within their tested settings. Uncertainty alone does not establish that a model is safe, and a confident prediction can still be wrong. I want to understand when these signals support useful intervention and how their reliability changes as models learn and act.
Long-term direction
Making capability useful for human oversight.
My long-term goal is for increasingly capable agents to help establish the conditions under which their actions can be trusted. This includes identifying requirements that need clarification and seeking evidence when a decision remains unresolved.
I am interested in whether training can make failures easier for independent supervisors to detect, and whether useful uncertainty signals remain reliable under changing incentives. The goal is to improve oversight while preserving the ability to complete tasks effectively.