Open benchmark · v0.1 RC

Can an LLM reviewer tell a real defect from a convincing false alarm?

WebFixBench is a reproducible benchmark for measuring the reliability of LLM-assisted code review on web-application changes — not just whether a model can produce a patch.

No LLM judgeHuman-reviewed ground truthApache-2.0
review-result.json
PrecisionMeasured
RecallMeasured
False alarmsMeasured
16cases in php-web-v0.1
12cases with a labelled defect
4clean controls
15frozen human-reviewed labels

The benchmark

Review quality is more than finding bugs.

A reviewer that flags everything can look impressive on a dataset containing only defects. WebFixBench deliberately measures the failure modes that make review tooling hard to trust.

01

Detection

True positives and false negatives are scored against explicit per-case labels.

02

False positives

Four tempting clean controls make indiscriminate vulnerability reporting cost something.

03

Structural reliability

Malformed responses are counted explicitly and never silently treated as correct.

04

Confidence

Per-finding and overall confidence are stored so confidence can be compared with correctness.

05

Cost & latency

Runs capture wall-clock latency and provider-reported usage, with optional user-supplied pricing.

06

Ecosystem breakdowns

Results can be broken down across PHP, Laravel, WordPress and the supported defect categories.

Current suite

PHP web ecosystem first.

The engine is language-agnostic, but v0.1 makes a deliberately narrow claim. The current suite contains synthetic PHP web changes designed to isolate one defect at a time and keep the ground truth decidable.

Read the dataset documentation ↗
EcosystemCasesShare
Laravel8
WordPress5
PHP3
authorization · 3injection · 4xss · 3secrets · 1unsafe deserialization · 1clean control · 4

Methodology

Ground truth is frozen before evaluation.

Cases are authored or accepted by a human maintainer. Model output is never used as the answer key. Matching is deterministic, versioned and reproducible.

1

Fixed input. A unified diff plus explicit context defines exactly what the reviewer may assume.

2

Human label. Expected findings are reviewed before model evaluation.

3

Deterministic scoring. Findings match by normalized defect type, optionally also by file — no LLM judge.

4

Raw runs preserved. Paid model outputs can be rescored later without making another API call.

Read the full methodology ↗

Results

No real-model leaderboard yet.

WebFixBench v0.1 is currently a release candidate with a working evaluation harness and a small labelled dataset. Real-model results have not been published, so the correct result is still not yet measured.

A committed mock-provider report exists only to exercise the pipeline; it is not a model benchmark result.

Published model resultsComing later
No published measurements

Reproduce it

Run the whole pipeline locally.

Python 3.9+, no runtime dependencies for the core package. Paid providers are opt-in and require an explicit model ID.

terminal
git clone https://github.com/lex127/webfixbench.git
cd webfixbench
pip install -e .

webfixbench run --provider mock --out results/my-run.json
webfixbench evaluate results/my-run.json
webfixbench report results/my-run.json --out results/my-run.md

Roadmap

Small by design. Broader only when the labels can support it.

Now · v0.1

Make the baseline publishable

Freeze the remaining case, complete academic review, validate live provider adapters and add a second reader.

Next · v0.2

Increase evidential value

Grow the PHP suite, add carefully sourced real-world cases, multi-finding cases, harder controls and inter-rater agreement.

Later

Expand context and ecosystems

Confidence calibration, context-level experiments, JavaScript/TypeScript and Python web suites.

Limitations

What v0.1 does not prove.

16 cases are small. Per-category numbers rest on only one to four cases.

The cases are synthetic. They isolate textbook patterns; they do not establish real-world review performance.

One maintainer reviewed the frozen labels. Inter-rater agreement is not yet measured.

This is not certification. A good score does not mean a model or tool is safe enough to replace human review.

People

Built for reproducible evaluation.

VC

Valeriia Chumak

Academic Collaborator · Evaluation Design & Research Framing

KhNURE profile ↗

Open source

Inspect the cases. Re-run the scoring. Challenge the labels.