Hypermetron
← All posts
#AI#Tooling

Prompts are code. So we built a linter for them.

Teams review every pull request, then ship the prompts that drive their AI features with no review at all. Rubrkit is our answer to that gap.

In most codebases we work on, a one-line change to a function goes through review, CI, and a test suite. A two-hundred-word change to the prompt behind the product’s AI feature goes through… whoever wrote it reading it once and deciding it sounds right.

That gap bothered us. Prompts, agent specs, skills, and multi-step workflows are instructions a system executes. They have inputs, expected outputs, and failure modes. They are code in everything but syntax — and they deserve the same engineering discipline.

Why prompts fail quietly

A broken function usually fails loudly: an exception, a red test. A weak prompt fails politely. The model still answers, the output still looks plausible, and the problem only shows up later as inconsistent results, edge cases handled differently on every run, or an agent that wanders off task.

When we looked at why instructions underperform, the causes were rarely exotic. They were the same handful of omissions, over and over:

  • Objective — the task is implied rather than stated.
  • Output specification — no format, length, or structure, so every run invents its own.
  • Context — the model is asked to decide things it has no information about.
  • Constraints — nothing says what is out of bounds.
  • Evaluation criteria and verification — no definition of a good answer, and no way to check one.

None of those are about clever wording. They are specification problems — the kind engineers already know how to catch.

An example

Here is the kind of instruction that ships every day:

Summarize this support ticket and suggest a fix.

It reads fine. But it doesn’t say who the summary is for, how long it should be, what to do when the ticket is missing information, whether “a fix” means a reply to the customer or an internal note, or how anyone would tell a good answer from a bad one. Each of those gaps is a place where two runs can disagree — and neither is technically wrong.

What Rubrkit does

Rubrkit treats an instruction the way a linter treats a source file. You paste in a prompt, agent spec, command, skill, or workflow, and it:

  • Scores it against a ten-dimension rubric — objective clarity, output specification, context, constraints, evaluation criteria, verification, and more.
  • Returns a critique that points at the weak dimensions, not just a number.
  • Hands back a rewritten version that closes those gaps.
  • Pairs the rewrite with simple evals, so you can check that the new version actually behaves better instead of taking our word for it.

That last step is the one we care about most. A rewrite that only reads better is an opinion. A rewrite that ships with the checks to prove it is an engineering change.

Where it fits

Rubrkit runs in the browser for one-off reviews. For teams, it also ships as a CLI (the rubrkit package on npm) and as an MCP server, so instruction reviews can sit inside the same editor and agent tooling where the instructions are written.

The goal isn’t a perfect score. It’s moving prompts out of the “read it once, looks fine” category and into the same loop as the rest of the codebase: written against a spec, reviewed against a standard, and verified before they ship.

You can try it free at rubrkit.com — no card required.