The University of Sydney

Yingjie Bai柏英杰

Computer science and financial economics.

I'm an Honours student working on data valuation, robot learning, and how AI systems evaluate and use information.

How I got here

Working on data valuation, I noticed that the value of the same data changes with the model and the task. The evaluation method is part of the answer.

That got me interested in how people are evaluated, too. At UCLA, I took Sociology of Crime to learn how rules and deterrence shape behavior, and started thinking about the implications for AI safety. Economics and sociology often give me ideas that I then try to test in machine learning.

More on evaluation and value

Selected papers

Full manuscripts and shorter workshop versions are linked below.

TMLR2026Accepted

When Does Data Value Reduce to Class Balance? A Coverage View of Per-Point Data Valuation

A coverage view of data valuation, studying when selection in the learner's representation is worth the extra cost.

Research details & versions

The full version has been accepted by Transactions on Machine Learning Research (TMLR). The short version, “Whose Geometry? When Learner-Relative Data Selection Beats Input-Space Selection for Efficient Fine-Tuning”, was accepted at the 2026 SIGKDD Workshop on Resource-Efficient Learning for Knowledge Discovery (RelKD 2026). This work studies when it is worth paying for data selection in the learner's representation geometry rather than a cheaper input-space geometry.

The paper gives a kernel-coverage view of efficient fine-tuning: learner-geometry selection helps only when representation non-locality, task non-interpolability in input space, and a binding budget all hold. Experiments across vision and LLM tasks show that a cheap input-learner misalignment diagnostic can predict when learner-relative selection is useful.

First page of the data valuation manuscriptFull manuscript
RelKDKDD · 2026Workshop · accepted

Risk-Aware Data Auditing for Resource-Efficient Learning under Distribution Shift

Applying risk decomposition from finance to data valuation under distribution shift.

Research details & versions

A short workshop version has been accepted at the 2026 SIGKDD Workshop on Resource-Efficient Learning for Knowledge Discovery (RelKD 2026). Data-CAPM is a risk-aware group data auditing framework for resource-efficient learning under distribution shift.

For each evaluation scenario, a group return is the utility drop from leaving that group out; across scenarios, Data-CAPM decomposes returns into mean return, beta, alpha, residual risk, and Alpha/Risk. This helps distinguish robust idiosyncratic value from shortcut-like systematic exposure, redundant utility, and volatile contribution.

First page of the accepted Data-CAPM workshop paperWorkshop paper
All publications & work in progress8 entries

Eight first-author works — three accepted, two under review or in submission, three in preparation. † marks the extended version of a workshop paper (same line of work, not two results).

  1. † When Does Data Value Reduce to Class Balance? A Coverage View of Per-Point Data Valuation (full version) — accepted, Transactions on Machine Learning Research (TMLR), 2026.
  2. Risk-Aware Data Auditing for Resource-Efficient Learning under Distribution Shift (Data-CAPM, short version) — accepted, KDD RelKD 2026 workshop.
  3. Whose Geometry? When Learner-Relative Data Selection Beats Input-Space Selection for Efficient Fine-Tuning (short version) — accepted, KDD RelKD 2026 workshop.
  4. A Verifier Validated on a Proxy Can Reverse in Deployment: The Signed Effect of Audit Certainty — in submission, NeurIPS 2026 workshop “Who Verifies the Agents?”.
  5. Failure Is Lossy: Wrongful Convictions in Shared Failure Banks for Autonomous Research — in submission, NeurIPS 2026 workshop “Verification in the Age of AI Scientists”.
  6. † The Memory of Failure Should Be Lossy: Wrongful Convictions in Shared Failure Banks for Autonomous Research (full version) — submitting to ICLR 2027.
  7. Reliability-Gated Latent Coverage for Demonstration Selection in Imitation Learning — manuscript in preparation.
  8. Market-Based Coordination for Multi-Arm Robotic Systems via Dynamic Spatiotemporal Pricing — manuscript in preparation.

Lumen Agora

An academic exchange platform I'm building.

I initiated Lumen Agora as an open-source, non-profit academic platform for concise research discovery, verified academic identity, and community-driven knowledge exchange.

The project responds to a problem in fast-moving fields such as AI and robotics: there are too many important papers for most people to read in full, yet ideas still need to travel across disciplines.

We are a six-person team working on an early version and gathering feedback from researchers.

More about the project
Format

A "Xiaohongshu + LinkedIn" inspired model: concise summaries help key ideas and results travel quickly, while original papers and authors remain the source of depth and credit.

Purpose

The platform aims to increase author visibility, support cross-field discussion, and create a better channel between opportunity providers and students or early researchers.

Stage

A six-person early team is exploring software, cybersecurity, and outreach while collecting feedback from researchers and faculty across fields.

Education & experience

Education

I am a final-year Bachelor of Advanced Computing (Honours) student at The University of Sydney, majoring in Computer Science and Financial Economics, with honours supervision by Dr. Weiming Zhi.

Curriculum vitae

Feb 2023 - Nov 2026

The University of Sydney

Bachelor of Advanced Computing (Honours), majoring in Computer Science and Financial Economics.

Honours supervisor: Dr. Weiming Zhi

Jun 2026 - Aug 2026

University of California, Los Angeles

Attended UCLA Summer Sessions through The University of Sydney's United States Studies Centre, taking COMM 109 - Entrepreneurial Communication and SOCIOL 147A - Sociology of Crime. The sociology course asked how deterrence, rules, and institutions shape behavior — and that question became a paper: my work on the signed effect of audit certainty for AI verifiers, now in submission at a NeurIPS 2026 workshop.

Sep 2025 - Dec 2025

National University of Singapore

Exchange student supported by the Vice-Chancellor's Global Mobility Scholarship, studying robotics, financial economics, game theory, and machine learning for data mining. During this period, I also joined Lin Shao's lab, where robotics coursework and research practice began to reinforce each other.

See related research experience

Jun 2021 - Nov 2022

The University of New South Wales

Completed foundation studies in science in Wuxi, China, before beginning undergraduate study in Australia.

Research experience

Aug 2025 - May 2026 National University of Singapore

Research Intern, Lin Shao Research Group

Worked with Chenrui Tie on complex dual-arm task planning and execution, and with Chongkai Gao on benchmarks for evaluating robotic manipulation systems.

This work connects high-level language reasoning with embodied action: how a robot decomposes goals, coordinates arms, executes long-horizon tasks, and remains reliable when the environment changes.

Robotics LLM planning Benchmarks
Nov 2024 - Feb 2025 Peking University

Research Intern, Songfang Huang Research Group

Built a fully localized deployment for a domain-specific energy AI system in collaboration with ENN Group, including local LLM inference, databases, multiple RAG frameworks, and an interaction interface.

The project gave me hands-on experience with domain knowledge retrieval, prompt grounding, private deployment, and evaluation for practical LLM systems beyond general-purpose chat settings.

RAG Energy AI LLM systems
Jul 2023 - Sep 2023 Tsinghua University

Research Intern, Kun Tang Research Group

Worked with Dr. Ziheng Meng using open satellite imagery, public datasets, and machine learning to study relationships between under-five child health, mortality, environment, medical infrastructure, and regional health disparities in Africa.

This experience shaped my interest in using large-scale data and computational methods to inform public policy, especially where inequality is difficult to observe directly.

Remote sensing Public health ML analysis

Academic service

  • NeurIPS 2026 Workshop AI4ScienceReviewer
  • NeurIPS 2026 Workshop Verify-AgentsReviewer

Leadership and engagement

What data valuation made me think about

A personal note

I don't think we can do without evaluation. But I increasingly think a score should come with an account of its assumptions: what it rewards, what it misses, and the conditions in which it holds. This is the thought that stayed with me after the paper.

Read the full reflection

What began as a technical question about data valuation became, for me, a broader question about evaluation itself: whether a metric is measuring value, or simply measuring adaptation to the current environment.

I now see evaluation systems as environments too. A score can be useful, but it should also reveal its assumptions, risks, blind spots, and the conditions under which its judgment holds.

I do not oppose evaluation. Without evaluation, societies cannot allocate resources and organizations cannot make decisions. But I increasingly believe that responsible evaluation should not only produce a score or ranking; it should also explain the environment in which that judgment is valid, the assumptions it depends on, what it rewards, what it ignores, and whether it identifies long-term value or merely short-term effectiveness within the current system.

This is the most important thought the paper left me with: evaluation systems are themselves environments. When a person is evaluated, they are not revealing value in a vacuum; they are being observed through a particular set of rules, resources, languages, and metrics.

Some people may appear more valuable because they learned earlier how to speak the system's preferred language, while others may have potential that current metrics cannot capture. Low scores do not always mean low value, and high scores do not always mean robust contribution.

My hope for future AI and robotics is cautiously optimistic: if technology can take over more repetitive, inefficient, and low-meaning work, perhaps people will have more room to explore different paths instead of being fixed too early by a single metric. In machine learning terms, I hope technology can expand each person's search space.

Work in progress

Some of these ideas have preliminary experiments; others are still at the conceptual stage.

Data valuation and selection

Data selection under limited budgets and distribution shift.

Research notes

How the value of data should be measured and allocated under budgets and distribution shift. Data-CAPM imports CAPM-style risk decomposition into group-level valuation; a coverage theory proves when per-point scores must fail and when selection should happen in the learner's own geometry.

Robot learning and demonstration selection

Selecting demonstrations for imitation learning, with a self-built SO-101 arm for experiments.

Research notes

Budgeted demonstration selection for imitation learning, reframed as set-level latent-mode coverage with reliability gating and evaluated in closed-loop rollouts — grounded in dual-arm planning experience at NUS and a self-built SO-101 arm on the LeRobot stack.

Verification and memory for AI agents

Auditing AI agents and revisiting ideas that were incorrectly recorded as failures.

Research notes

The economics of oversight for agentic systems: when raising audit certainty deters and when it backfires (a signed certainty law), and why the shared memory of failures in autonomous research should be lossy — wrongful convictions lock out good ideas.

Research approach & motivation

One question organizes my research: how should learning systems value, select, and trust information? My method is transferring mechanisms that other fields have already stress-tested — finance, the economics of crime, legal institutions — into machine learning, carried over as theorems and experiments rather than metaphors, together with a characterization of when the transferred mechanism fails.

Having studied across different educational and cultural environments, I became increasingly aware of how unevenly opportunity is distributed and how strongly access to knowledge and guidance shapes individual trajectories. I hope to explore how intelligent systems and knowledge infrastructure can lower barriers to knowledge and opportunity, making technological progress more broadly accessible in an increasingly AI-driven world.

Multi-arm coordination and resource mechanisms

Beyond my current papers, I am developing a set of connected ideas around how multi-arm systems can coordinate under local information, shared constraints, limited resources, and dynamic task value.

03

Multi-arm coordination through pricing and compressed information

In multi-arm environments, trajectory planning and task allocation become increasingly complex as the number of arms grows. A fully centralized planner can quickly face high computational cost, limited scalability, and weak real-time performance.

My intuition comes from resource coordination in human society: prices and money compress information about scarcity, preference, and constraints without transmitting every detail. Inspired by this, I am exploring dynamic prices for time, space, path conflicts, and shared resources, while keeping collision avoidance as a hard constraint.

With limited personal compute, I have mainly tested this idea in simplified two-dimensional environments. The early results suggest feasibility, but scaling it to full multi-arm simulation or real robotic systems would require more compute and experimental support.

04

Arm quantity, task scale, and survival cost

I am also interested in a higher-level question: for a given task scale, how many arms should a system activate? More arms do not necessarily produce linear efficiency gains; they can also create more collision constraints, computation costs, maintenance burden, depreciation, and competition for space.

I therefore consider giving each arm an activation, operating, or survival cost, including electricity, hardware depreciation, GPU resources, maintenance, downtime, and the opportunity cost of occupying space and time. A system could then activate only the arms whose marginal contribution justifies their cost.

This idea is partly inspired by resource economics: using a scarce resource now can impose opportunity costs on future tasks. In multi-arm systems, space, computation, and machine lifetime can also be treated as finite resources rather than free inputs.

05

Dynamic coalition and rental-style collaboration

Some manipulation tasks, such as furniture assembly, large-object transport, or complex assembly, require temporary collaboration among multiple arms. A single arm may be unable to complete a step without calling nearby arms into the task.

I am exploring a dynamic coalition mechanism in which PDDL or other task-planning methods decompose a complex task into subtasks with assigned rewards. When a subtask requires multiple arms, its reward can increase, giving suitable arms an incentive to join temporarily.

This creates a rental-style collaboration pattern: an arm can request assistance for a high-value time window; other arms weigh their own tasks, motion costs, opportunity costs, and collaboration reward; after the subtask is completed, each arm returns to independent execution.

Technical projects

Hardware

Self-Built SO-101 Robot Arm (LeRobot)

Assembled and calibrated a low-cost open-source manipulator and run it on the LeRobot stack for teleoperated demonstration collection — a personal hardware testbed for the demonstration-selection research.

Robotics

Multi-robot coordination via pricing mechanisms

Proposed and implemented a framework where multiple robots coordinate through dynamic pricing of shared spatiotemporal resources, exploring market-inspired mechanisms for decentralized task allocation and congestion management.

I am especially interested in whether ideas from economics, such as pricing and resource allocation, can help decentralized robots negotiate congestion and shared constraints.

My intuition is that multi-arm coordination in constrained space is partly a resource allocation problem: each region of space has time-dependent scarcity, and price can act as both a coordination tool and an information signal.

Planning

Robot Task and Motion Planning Project

Developed a robotic task organization system integrating large language models, PDDL planning, and agent frameworks to study multi-robot coordination and complex task planning.

Financial ML

Loan Prediction

Developed predictive models using GNN and CatBoost with engineered financial and online behavioral features to estimate loan levels.

Data Analysis

Urban Traffic Data Analysis

Analyzed open transport data for the Sydney Light Rail using R for time-series modeling and visualization, studying passenger-flow dynamics and implications for congestion management.

Technical skills and interests

Technical

Python, C, Java, R, Git, Linux (self-hosted Raspberry Pi server), machine learning, econometrics, and data analysis.

Robotics and AI

Multi-agent systems, robotics, LLM agents, AI safety and agent verification, RAG systems, planning, and simulation — plus a self-built SO-101 arm on the LeRobot stack.

Interests

Fencing, kayaking, international travel, game theory, philosophy of algorithms, technology and society.

Contact

Email is the easiest way to reach me.