Any model. Your keys. No markup.
Paste an API key, or connect your claude.ai account by device authorisation — you never see a terminal. Then decide which model does which job: Opus for architecture, Sonnet for standard work, a local Qwen for boilerplate, Haiku for docs. Providers bill you directly.
| Role | Chain | State |
|---|---|---|
| architecture | opus → gpt → glm | Ready |
| standard | sonnet → glm-flash | Ready |
| boilerplate | qwen-local → haiku | Ready |
| tests | sonnet → qwen-local | Cost cap reached |
| docs | haiku → gemini | Needs an API key |
Anthropic · OpenAI · Google · Z.ai · local
API key or device flow. Local means your own inference host — the same way we run a Qwen model on our own GPU for cheap work.
Per-role ordered chains
Each role gets an ordered list of models. Failover on provider error or when a cost threshold is hit. A cost view shows what each role spends.
served_by on every reply
The CEO reports which model actually answered. Failover is visible in the reply, never silent.
Catalation never bills AI.
A platform that resells tokens has an incentive to make your agents chatty. We don't have that incentive. You pay your providers at their price; we charge a flat subscription for the company around them. Keys are stored through an envelope-encrypted secrets broker and never appear in the company document, a log or a response.
How keys are protected“A router picks the right model per task. Catalation never touches AI billing.”