CampStudio AI Lab · Experiment 01
The Lazy Vibe Coder Benchmark
How much useful product work can an AI coding model deliver from one deliberately underspecified prompt?
- Output
- Working service
- Evaluation
- Human-scored / 100
- Sample
- 1 run per setup*
Results
Useful work versus resource use
Method
One vague brief. One clean agent. One shot.
A fresh OpenClaw agent receives a deliberately non-detailed product brief. Its job is to turn that single prompt into a finished, usable service without clarification rounds or steering.
-
01
Fresh setup
A clean OpenClaw instance starts with the selected model and reasoning effort.
-
02
Underspecified brief
The same intentionally loose task leaves room for judgment, initiative, and product sense.
-
03
One-shot build
The agent must produce the service in one pass, not improve it through a back-and-forth conversation.
-
04
Human review
The result gets a composite score out of 100 for brief fit, integrity, design quality, and useful features nobody explicitly requested.
Real-world efficiency
Credits, not token-list prices
Cost efficiency is measured in credits consumed from the PRO x20 plan. Credits are quota units, not dollars and not raw tokens: the more credits a run uses, the more of the plan allowance it burns.
That matters because a model's published price per million tokens is not always proportional to its real impact on a subscription quota. Measuring score per credit compares the work you receive with the allowance you actually lose.
Raw results