CampStudio AI Lab · Experiment 01

The Lazy Vibe Coder Benchmark

How much useful product work can an AI coding model deliver from one deliberately underspecified prompt?

Output
Working service
Evaluation
Human-scored / 100
Sample
1 run per setup*

Results

Useful work versus resource use

Score / 100 against total PRO x20 credits consumed. Left and up is better. Lines connect Low to Max reasoning. Only points are interactive.

Method

One vague brief. One clean agent. One shot.

A fresh OpenClaw agent receives a deliberately non-detailed product brief. Its job is to turn that single prompt into a finished, usable service without clarification rounds or steering.

  1. 01

    Fresh setup

    A clean OpenClaw instance starts with the selected model and reasoning effort.

  2. 02

    Underspecified brief

    The same intentionally loose task leaves room for judgment, initiative, and product sense.

  3. 03

    One-shot build

    The agent must produce the service in one pass, not improve it through a back-and-forth conversation.

  4. 04

    Human review

    The result gets a composite score out of 100 for brief fit, integrity, design quality, and useful features nobody explicitly requested.

Real-world efficiency

Credits, not token-list prices

Cost efficiency is measured in credits consumed from the PRO x20 plan. Credits are quota units, not dollars and not raw tokens: the more credits a run uses, the more of the plan allowance it burns.

That matters because a model's published price per million tokens is not always proportional to its real impact on a subscription quota. Measuring score per credit compares the work you receive with the allowance you actually lose.

Raw results

Every tested configuration