Ponytail 5 · rebuilt from the ground up

Half the code.
And yet better.

A ruleset that makes your AI coding agent write the least code that works: 53% less code, 41% faster, 26% cheaper, and just as correct.

01Benchmark

Same agent. Same tasks.
Less of everything.

Opus 5.5 on 39 tasks, five runs each: feature tickets in a real repo, tasks with hidden checks, open "build me" requests. No skill = 100%.

-53%lines of code
100%
52%
47%
-45%output tokens
100%
57%
55%
-26%cost
100%
84%
74%
-41%time
100%
62%
59%
{…}no skill{…}Ponytail v4.13{…}Ponytail 5

Geometric mean of per-task medians. Hidden correctness and safety checks: 96% / 96% / 97% passed.
Validation, error handling, security and accessibility are never simplified away. Method and limits.

02Tests

and yet

98%of risky logic ships with a test. A branch, a loop, a parser, money or security: one small test, right where it counts.
Without Ponytail68%
03Before / after

Your agent reaches for a library.
Ponytail reaches for the browser.

"Add a date picker to the frontend." Same agent, same repo, lines added, median of five runs.

335lines without Ponytail

Without Ponytail: a hand-built calendar and date picker, 335 lines. With Ponytail 5: one 10-line file, built on the browser's own <input type="date">.

The repo is React, so the answer stays React. The calendar comes from the browser.

04How it works

Stop at the first rung
that holds.

It reads the code first and traces the real flow. Then it climbs. Pick a request:

  1. 01Does this need to exist?no: skip it
  2. 02Is it already in this codebase?reuse it
  3. 03Does the standard library do it?use it
  4. 04Does the platform have it?use it
  5. 05Does an installed dependency solve it?use it
  6. 06Can it be one line?one line
  7. 07Only then: the minimum that works.plus one test if it has logic

05Review

The review, rebuilt.
Now it finds what the diff hides.

The old review only looked for code to cut. The new /ponytail-review reads the code your change touches, not just the diff: bugs, security, real load, missing tests, speed, and what to cut. Each finding says what the code does, what goes wrong, how to fix it, and what happens if you don't. /ponytail-audit does the same for the whole repo.

›

    100%of planted problems foundWithout Ponytail 87%
    100%of problems outside the diff foundWithout Ponytail 78%
    06New in 5

    Rebuilt,
    not patched.

    rules = rules.filter(beats_baseline)

    Half the rules.

    Every line had to beat the version without it, one change at a time. What did not win, left.

    assert parse_price("1,50") == 150

    Tests where it counts.

    A branch, a loop, a parser, money or security leaves one small test. Trivial stays trivial.

    // skipped: retries, add when it flakes

    It tells you what it skipped.

    Every reply ends with what was left out or not checked, and any risk.

    claude · codex · copilot · gemini · pi

    Works where you work.

    Claude Code, Codex, Copilot, Gemini, Cursor, OpenCode, pi and a dozen more.

    07Install

    Two lines.
    Then it's always on.

    /plugin marketplace add DietrichGebert/ponytail
    /plugin install ponytail@ponytail

    Cursor, Windsurf, Cline, Kiro, OpenCode, Zed and the rest: INSTALL.md.
    Only install from DietrichGebert/ponytail or @dietrichgebert/ponytail on npm.

    /ponytail lite|full|ultra|off
    how lazy: full is the default
    /ponytail-review
    a full review of the diff
    /ponytail-audit
    the same for the whole repo
    /ponytail-debt
    deferred shortcuts as a ledger
    /ponytail-gain
    the benchmark scoreboard
    /ponytail-help
    quick reference