Virtual research collective · Distributed

Publication ·

New Review clarifies how to validate LLMs used as human proxies

A Nature Computational Science Review by Nikita Karetnikov, Iyad Rahwan, and Davor Svetinovic separates four uses of LLMs as human proxies and matches each to a different validity question.

Nature Computational Science has published “Large language models as human proxies,” a Review examining how large language models are used to stand in for people across scientific and applied settings.

The paper distinguishes four roles: believable agents, task agents, experimental subjects, and silicon samples. Its central methodological point is that “human-like” is too broad to serve as an evaluation standard. Each role supports a different kind of claim and therefore needs a different human reference, test, and validity criterion.

Nikita Karetnikov and Davor Svetinovic are represented in the REQS community through Trust Lab, and the journal record lists REQS Labs among Svetinovic’s affiliations. Iyad Rahwan contributed as a co-author. This announcement records that connection without treating the collective as the sole producer of the work.

REQS Labs especially congratulates Nikita Karetnikov, a Trust Lab researcher, for his leading role in surveying the literature and writing the paper. The work also reflects the collective’s wider standard for agentic AI research: define the claim, specify the model’s role, and choose evidence that can actually support the claim.

What the Review changes

One phrase, four different claims.

01

Four distinct roles

Believable agents, task agents, experimental subjects, and silicon samples are separated rather than grouped under a single claim of human likeness.

02

Role-specific evidence

The relevant comparison and test depend on whether the claim concerns believability, task competence, machine behaviour, or human populations.

03

Clearer research claims

The framework helps researchers avoid moving from one kind of similarity to conclusions that require a different form of validation.