Train a small model on your data.Own the weights. Run it anywhere.
Bring your data, or just describe the task. Get back a small specialized model, one native artifact that runs anywhere, with weights that are yours to keep.
Dead simple
Upload data, pick a base, train. No ML knowledge required, full control is one click deeper.
Runs anywhere
The same native artifact runs on-device, in Node, or hosted. Try it in your browser the moment it's done.
Yours to keep
Download the weights and run them forever, wherever. Nothing about your model is locked in.
Replace your frontier API bill
You're paying per token to a giant general model to do one narrow, repetitive job. Point Gerbil at your production logs and it distills that traffic into a specialist that matches quality at a fraction of the cost, runs on your own hardware, and keeps your data in-house.
- Distill directly from your existing API logs, no labeling
- Same task quality, roughly 50x cheaper per call
- Runs on your infra or on-device, data never leaves
- Tune a tiny specialist or a denser general model, your call
Tune a model
Start a job, watch it train, and try the result on-device.
You know your task. We do the ML.
Show us examples of the job done right, and Gerbil handles the GPUs, the training, and the export for you, so your team goes from raw data to a model you own without hiring an ML engineer or touching infrastructure.
Bring data, or describe it
Upload a JSONL/CSV, or just describe the task and Gerbil generates a dataset for you.
Train
We fine-tune your chosen base on managed GPUs. You never touch infra, just watch it run.
Try it
The moment it's done, chat with your model right in the browser, side-by-side with the base.
Own it
Download the native weights or one-click host them. Runs in the browser, Node, or your server.
A specialist beats a generalist at one job
A tuned small model matches frontier quality on your task at a fraction of the latency, size, and cost, running on your own hardware. Same accuracy, milliseconds instead of seconds, and nothing per call.
| Task | Frontier API | Gerbil tuned | Latency | Size | Cost / 1K calls |
|---|---|---|---|---|---|
Support triage classify | 93% | 94% | 41ms | 400MB | $0.00 |
Invoice to JSON extract | 95% | 95% | 44ms | 400MB | $0.00 |
Text to SQL generate | 91% | 90% | 47ms | 400MB | $0.00 |
Tone rewrite rewrite | 92% | 93% | 39ms | 400MB | $0.00 |
Narrow specialist, or your whole org's model
Train a tiny model for one job (faster, cheaper, on-device), or a denser general model that carries your org's knowledge in its weights. Same flow, your call.
Classify & label
Route support tickets, moderate content, tag records.
“my card was charged twice” → billingExtract structured data
Pull clean JSON fields out of messy text.
invoice text → { total, date, vendor }Rewrite & format
Enforce a house tone or an exact output shape.
draft → on-brand replyRoute & triage
Send each request to the right queue or tool.
request → { team, priority }Translate & localize
Domain-specific translation your terms depend on.
EN spec → localized copySummarize to a schema
Turn long text into a fixed, parseable summary.
thread → { decision, owners }Cache the base once. Swap flavors on top.
A tune is a small delta on a base model. Keep the heavy base cached once and load each specialization as a flavor , megabytes, not gigabytes. Snap on a support agent, hot-swap to a refunds classifier, all on the same cached model.
- Tiny per-flavor updates ship over-the-air, the base already lives on the device
- Runs offline once the base and a flavor are cached, no network at inference
- Stays private: prompt, flavor, and output never leave the device
- One base, a library of flavors, swap skills instantly, no model per task
await g.loadModel("qwen3.5-0.8b");
// heavy base — cached once
await g.loadAdapter("hf:eyal/support-flavor");
// MB-size flavor on top
// swap skills — the base never re-downloads
await g.loadAdapter("hf:eyal/refunds-flavor");Coming , an adapter marketplace to share and install flavors (the next “skill”), and multi-flavor hosting to serve many flavors per GPU off one loaded base.
Set it once. It compounds while you sleep
A one-time model decays as your world changes. Hosted +Autotrain connects your live data and keeps the model sharp on its own, every version is the same native artifact, promoted only if it's genuinely better. This is the tier that keeps paying you back.
Task quality climbs with each cycle and holds, you capture the gains without lifting a finger.
Source connectors
GitHub, Linear, Notion, Slack, or a webhook. New examples flow in as your product runs.
A/B canary
Every retrain is tested on a slice of live traffic before it ever reaches everyone.
Auto promote / rollback
Better gets promoted, worse gets rolled back automatically. You're never exposed to a regression.
Version history
Every version is kept and diffable. One-click rollback to any prior model, anytime.
The weights are yours, and they run everywhere
Own your weights
Download the model and keep it forever. No subscription required to keep using what you trained.
Runs anywhere
One native artifact runs in the browser, in Node, on your server, or on-device, no separate builds.
No conversion step
The engine loads your tuned weights directly. Train, then run, nothing in between.
Private by default
Your data trains your model and is yours. Run inference on-device and it never leaves the machine.
Pay to own, not to try
The demo is free, so you see it work before you pay. Then it's one flat fee to train on your data and own the model, and expansion from there.
Try the whole flow in your browser. See it work first.
- Tune a sample task on-device, no account
- In-browser playground, tuned vs. base
- See the accuracy lift before you pay
Train on your data and own the native weights.
- Real training on managed GPUs, your dataset
- Starts at $99 (up to 1,000 examples), scales per-example
- Download the model, yours to keep, no lock-in
A private endpoint, live in seconds.
- 1M inference requests / mo included
- Then $0.20 / 1K requests, scales to zero when idle
- Swap versions without code changes
Hosted, and it retrains itself on a rolling window.
- Monthly $499, weekly $1,499, or daily $4,999, per model
- Hosting included; fresh weights to download each cycle
- A/B canary + auto-rollback; version history
Introductory pricing, final tiers may change.
Questions
Do I actually own the model?+
Yes. Export downloads the native weights and they're yours to run anywhere, forever, no subscription needed to keep using them.
What data do I need?+
A JSONL/CSV with at least ~100 examples (chat messages, or prompt/completion pairs). No data? Describe the task and Gerbil generates a quality-filtered dataset for you.
Where does it run?+
The tuned model is one native artifact that runs on the Gerbil engine in the browser, in Node, on your own server, or on-device. The same weights, everywhere.
Which base models can I tune?+
Qwen 3.5 (0.8B and 2B) today, with LFM2.5 and Gemma 4 being added, all families the engine already runs natively.
How long does training take?+
Small models on a focused dataset typically finish in minutes. You watch the lifecycle live and try the result the moment it's done.
What does it cost?+
The in-browser demo is free, so you can try the whole flow first. Training on your own data and owning the weights starts at $99 per model, that covers up to 1,000 examples and scales per-example with your dataset size, and it includes downloading them (no separate export fee, the model is yours). Synthetic data adds a small per-example generation fee. From there it is optional: host it for $49/mo, or host and auto-tune it from $499/mo.
Can I try it before I pay?+
Yes. The in-browser demo runs the entire flow, pick a base, add data, train, and chat with the result, for free and with no account. You only pay once you want to keep the model.
What is host and auto-tune?+
Hosting ($49/mo) gives your model a private endpoint with 1M requests a month included. Host and auto-tune (from $499/mo) adds it: the model retrains itself on a rolling window of your new data on a set cadence, A/B canaries each new version, and only promotes it if it is genuinely better. Cadence tiers are $499/mo monthly, $1,499/mo weekly, and $4,999/mo daily.
Can I add more training data automatically?+
Yes. If your dataset is thin, +Synth has a frontier teacher generate more examples in your style, quality-filtered and deduped against what you have. Pick exactly how many with a slider; it folds into your training order, so you pay once for precisely the amount you choose.
How do I turn my production calls into training data?+
Capture. Every call you already make to a frontier model is a labeled example you paid for. Tag each call-site by task and each tag becomes a clean dataset you can train a specialist from, instead of one firehose that only yields a generalized model. Capture docs →
Do you offer dedicated hosting or enterprise (SLA, VPC, SOC 2, HIPAA, on-prem)?+
Yes. Dedicated hosting, SLAs, VPC isolation, SOC 2 / HIPAA, and on-prem deployments are all available. Email sales@gerbilsdk.com and we'll size it with you.
Train your first model
It's free to train and try. Bring data or describe the task, you'll have a model you can run in your browser in minutes.