• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
AI
Report

ScientistTwo is claimed to run experiments and write papers after a researcher sets the problem

A post says it beat published human results on 86 of 107 problems drawn from accepted AI conference papers.

Christian SzegedyCS
Alex VeremeyenkoAV
2 Sources, 10h ago, first seen 10h ago

TLDR

A post says Google researchers built ScientistTwo to form hypotheses, run experiments, write papers and simulate peer review after a human sets the problem. It reportedly beat human results on 86 of 107 problems from accepted ICLR, ICML and NeurIPS papers. The post cites a 25.2% average gain but a 7.7% median. Nine reviewers reportedly rated 33 AI papers tied with accepted human papers overall, with a small methodological rigor edge for the human papers.

Combined views

5.7K

2 Sources, first seen 10h ago

65 likes5 comments59 saves28 reposts

Combined views

5.7K

2 Sources, first seen 10h ago

65 likes5 comments59 saves28 reposts

Sentiment

Positive——Negative

Summary

Not enough discussion yet.

No sentiment analysis available yet.

Featured Source

Sentiment

Positive——Negative

Summary

Not enough discussion yet.

No sentiment analysis available yet.

2 Sources

Alex Veremeyenko@alex_veremDid Google just automate the PhD? Google researchers built an AI that reads a problem, forms hypotheses, runs experiments, writes the paper, and simulates its own peer review. No human steps in once the problem is set. It's called ScientistTwo. They gave it 107 problems from papers accepted at ICLR, ICML, and NeurIPS, and it beat the published human result on 86 of them. The average gain over the human method was 25.2%. The average paper took about 2.5 days and cost about $3,765 in model and compute costs. Nine human reviewers compared 33 of these AI papers with accepted human papers. Their overall verdict was a tie. The same paper has the less exciting numbers too. The reviewers gave the human papers a small edge on methodological rigor. The authors say it isn't producing spotlight-level work yet. The median gain was 7.7%, so a few big wins pull the average up. Without its checking agents, 19 of its 1,840 citations were fake, and one solution broke the task rules to get a better score. With the checks on, both dropped to zero. A researcher still has to pick the right question. The AI cuts the wait for the answer to a few days. Would you trust a paper if you knew an AI ran every experiment in it? Paper's in the first reply 👇10h
Christian Szegedy@ChrSzegedyRT @alex_verem: Did Google just automate the PhD? Google researchers built an AI that reads a problem, forms hypotheses, runs experiments,…1h
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI

    2 Sources

    Alex Veremeyenko@alex_veremDid Google just automate the PhD? Google researchers built an AI that reads a problem, forms hypotheses, runs experiments, writes the paper, and simulates its own peer review. No human steps in once the problem is set. It's called ScientistTwo. They gave it 107 problems from papers accepted at ICLR, ICML, and NeurIPS, and it beat the published human result on 86 of them. The average gain over the human method was 25.2%. The average paper took about 2.5 days and cost about $3,765 in model and compute costs. Nine human reviewers compared 33 of these AI papers with accepted human papers. Their overall verdict was a tie. The same paper has the less exciting numbers too. The reviewers gave the human papers a small edge on methodological rigor. The authors say it isn't producing spotlight-level work yet. The median gain was 7.7%, so a few big wins pull the average up. Without its checking agents, 19 of its 1,840 citations were fake, and one solution broke the task rules to get a better score. With the checks on, both dropped to zero. A researcher still has to pick the right question. The AI cuts the wait for the answer to a few days. Would you trust a paper if you knew an AI ran every experiment in it? Paper's in the first reply 👇10h
    Christian Szegedy@ChrSzegedyRT @alex_verem: Did Google just automate the PhD? Google researchers built an AI that reads a problem, forms hypotheses, runs experiments,…1h
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet