Argentis Labs

Evaluation-first ML for decisions in regulated domains: lending, legal, insurance.

We build systems where a wrong answer has a cost: constrained decision models, human-in-the-loop agents, and the benchmarks that hold them accountable. Benchmarks are frozen and hashed before any model is trained against them.