Evalt

About

An open-source evaluation and routing project.

Evalt is built and operated by Jonathan Larson. It exists to make task-specific AI quality, cost, and routing decisions inspectable instead of relying on generic benchmark rankings.

Current status: beta · free early access · bring your own provider key · no managed paid plan is currently for sale

What Evalt does

Evalt freezes a quality contract for one recurring AI task, measures prompt/model/reasoning candidates under a provider-spend cap, and installs the lowest-cost configuration that clears the approved gates. The local SDK is the production runner; the optional hosted workspace receives privacy-bounded aggregate route metadata.

What it does not claim

A perfect observed score is not proof that unknown future inputs cannot fail. AI-generated cases are not human ground truth. A recorded example is not a universal model ranking. Evalt labels these boundaries and recommends version pinning, independent result review, and noncritical trials before production adoption.

Project and support

Source code, issues, and release history are public on GitHub. Use the issue tracker for support and non-sensitive feedback. Do not place credentials, customer content, or private vulnerability details in a public issue.