Send every prompt to the cheapest model that can answer it.
Sluice reads each request, picks the least expensive model that can handle it, checks the answer, and steps up a tier only when it has to.
Top modelanthropic/claude-fable-5.1
$0.015210
Routedanthropic/claude-haiku-5.5
$0.000152
You save $0.015058 (99.0%) on this prompt.
Estimate at list price for ~21 input and 300 output tokens. Classified as easy. Run it for real in the playground to see actual costs.
Two-line migration
Keep your OpenAI client. Change the base URL and the key.
TypeScript
const client = new OpenAI({
baseURL: "https://sluice.forum/api/v1", // was https://api.openai.com/v1
apiKey: process.env.SLUICE_KEY, // model: "auto" lets Sluice pick
});Create a key with a spending limit in the dashboard, under Gates.
How routing works
- 1Classify
Rules read length, code, math and the kind of ask, and sort the request into easy, medium or hard. No model call.
- 2Pick
The cheapest model in that tier that exists in the live catalogue and fits the request.
- 3Check
Empty, refused, cut-off or broken-JSON answers are retried once on the next tier up.
Prompts are never stored.