On 2 October, Constructor Labs hosted the Constructor Talks with Prof. Dr. Andrey Ustyuzhanin (Principal Investigator at Constructor Knowledge Labs, Adjunct Professor of Computer Science at Constructor University), on the current state of autonomous research agents and the challenges involved in assessing the quality of the research they produce https://lnkd.in/duJpMH_C The event attracted 140 participants.
The talks drew on an analysis of 139 public repositories and around 100 papers on AI scientist and AI researcher systems. Many of these systems can already generate hypotheses, code and scientific text, but evaluation remains a weak point.
Several issues appeared repeatedly across the systems reviewed:
– evaluation procedures were often weak or easy to modify;
– calibration data was missing or incomplete;
– published descriptions did not always match the code available;
– LLM-based judges were sometimes used without enough independent checks.
Prof. Dr. Andrey Ustyuzhanin argued that reliable research agents need evaluation methods that are tied to the domain and grounded in verifiable evidence. Different fields work with different types of data, methods and standards, so the same evaluation setup cannot simply be applied everywhere.
He also presented work being developed at Constructor Labs, where research tasks are structured around a defined outcome and supported by tools for simulation, fitting, scoring and verification. In one materials-science example, a domain-specific optimisation loop reduced the fitting error for a zinc-oxide dataset from roughly 80% to 10% without manual expert tuning.
The Q&A continued the discussion around verification, experimental validation, conflicting conclusions from different agents and the role of researchers in deciding what counts as sufficient evidence.
Key takeaways:
– Autonomous research agents are already capable of handling substantial parts of the research workflow, but reliable evaluation still lags behind generation.
– Verification needs to be built around the domain, with clear checks against measurements, simulations or other independent evidence.
– Research agents are currently stronger at searching and optimising within an existing framework than at changing the way a problem is framed.
– As these systems take on more research tasks, defining the problem and the criteria for a valid result becomes an increasingly important part of the researcher’s role.
Stay tuned for upcoming Constructor Talks, with new sessions and speakers to be announced soon.


