This is a Plain English Papers summary of a research paper called or follow me on is a new benchmark that aims to address this gap. It is a dynamic test that automatically generates and regularly updates a set of 1,000 questions about future events with no known answers at the time of submission. This ensures there is no risk of data leakage, which could artificially inflate a system's performance.
The researchers tested the forecasting capabilities of expert human forecasters, the general public, and large language models (LLMs) on a random subset of 200 questions from the benchmark. While LLMs have shown super-human performance on many tasks, the results here were different. The expert human forecasters outperformed the top-performing LLM in a statistically significant way (p-value = 0.01).
The researchers make the results publicly available on a leaderboard at or following me on Twitter for more AI and machine learning content.
SOCIAL SHARE CARD GENERATOR