LiveCodeBench
LiveCodeBench is a date-stamped benchmark for AI coding models, testing code generation, repair, execution reasoning, and performance on new problems.
LiveCodeBench is a continuously updated benchmark for evaluating how well AI language models solve programming problems. It tests whether a model can produce working code, repair mistakes, and reason about program behavior. Its distinguishing feature is time: problems carry publication dates, allowing evaluators to test models on challenges released after their training cutoff—the latest date covered by their training material. This helps separate genuine problem-solving ability from recalling previously encountered solutions.
Origins and Benchmark Design#
Introduced in the 2024 paper LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code, LiveCodeBench was created by researchers at UC Berkeley, MIT, and Cornell University led by Naman Jain. It collects competitive-programming problems from LeetCode, AtCoder, and Codeforces. Early headline baselines included GPT-4 Turbo and Claude 3 Opus, which performed strongly across its evaluation tasks; DeepSeek-Instruct-33B and Phind-34B were leading open-access baselines in the initial comparisons. These are historical reference points, rather than permanent leaderboard leaders.
The initial version, release_v1, contained 400 problems published from May 2023 through March 2024. Subsequent releases on Hugging Face expanded the collection, so LiveCodeBench has no single permanent size. A reproducible result must identify its release and evaluation window.
Like other benchmark datasets, it provides shared tasks and grading procedures for comparing models. Its date-stamped evaluation design additionally lets evaluators select a common period of newly published problems.
What LiveCodeBench Measures#
LiveCodeBench evaluates four related capabilities of large language models, systems that generate and interpret text, including source code:
- Code generation: Produce a program from a problem statement and example input-output pairs. The evaluator runs the solution against additional tests.
- Self-repair: Revise an incorrect solution after receiving execution feedback, such as an exception or a failing test.
- Code execution: Predict the output of a supplied program for a specified input. Here, “execution” describes the model’s reasoning task; the evaluator checks its prediction against actual execution.
- Test output prediction: Infer the expected output from the problem specification and a supplied input, without receiving the implementation.
The distinction between the last two tasks matters. Code execution asks what an existing program does; test output prediction asks what a correct implementation should do. Self-repair examines the feedback-and-revision behavior also used in agentic coding, where AI systems plan, write, run, and revise code.
For code generation, functional correctness means passing every required test. A program can run successfully yet produce incorrect answers, while an execution error can prevent it from producing an answer at all. Both failures affect benchmark performance.
Reading Scores and Understanding Contamination#
Pass@1 measures the likelihood that one generated solution passes all required tests, averaged across problems. A reported value of 60% means approximately six in ten problems are solved under the specified evaluation procedure. It does not mean that each program passes 60% of its tests.
Fair comparisons require matching the dataset release, date window, task, and generation settings. Prompt engineering, sampling settings, output limits, and repair opportunities can change the result. Easy, medium, and hard breakdowns also reveal differences hidden by an overall average.
Benchmark contamination occurs when evaluation problems or solutions appear in a model’s training data. This form of data leakage can inflate scores because the model has already encountered the material. Selecting problems published after the model’s cutoff reduces this risk, although the cutoff’s reliability and later model updates still matter.
A claim that a particular model “leads LiveCodeBench” therefore needs a dated leaderboard snapshot and matching evaluation settings. A ranking without that context is difficult to interpret.
Applications and Related Benchmarks#
LiveCodeBench supports comparing coding models and diagnosing whether weaknesses lie in generating solutions, understanding code, or repairing failures. Its contest-based tasks provide evidence about algorithmic programming ability; broader software-development evaluation also needs tests of repository navigation, dependency management, and integration behavior.
LiveBench is a separate, broader benchmark covering areas such as reasoning, mathematics, language, and coding. Similar names do not imply interchangeable scores.
For developers building with Ultralytics YOLO, LiveCodeBench offers one signal when evaluating a coding assistant. A concrete complementary workflow is using Ultralytics Agent Skills, which provide compatible coding agents with instructions for dataset preparation, training, inference, and export.
When that assistant helps build a YOLO26 application, evaluate its generated code separately from the vision model: validation mode measures prediction quality, while benchmark mode compares deployment-format accuracy and inference speed.










