Frontier-class results, on your own hardware.
The open models at the frontier now hold their own against the closed ones on the benchmarks your team actually cares about. You get that capability without the API, the vendor, or the exposure.
The open frontier, on your own hardware
The newest open models now match the big closed vendors on coding work, at a fraction of the cost. Shakudo runs them on your hardware, and your data never leaves your network. See the full roster, benchmarked side by side below.
Please enter your work email address.
By submitting your email you agree to receive occasional updates from Shakudo. See our Privacy Policy.


The open models at the frontier now hold their own against the closed ones on the benchmarks your team actually cares about. You get that capability without the API, the vendor, or the exposure.
Open-weight models are priced in compute, not per token. Run them on hardware you already own and the per-1,000-token rate stops climbing with every request you make.
Prompts and outputs never leave the network. No training on your data, no retention on their side, no foreign legal framework you have no control over.
Open weights mean you can steer the model. Point it at your documents, your tone, your domain, and it gets better at your work instead of guessing at everyone else's.
Shakudo fronts the whole roster behind a single OpenAI-compatible endpoint. Swap models as the frontier moves, and route each task to the one that does it best.
Closed-model capability plateaus and then prices itself up. The open line is still rising, and it does not require a seat, a contract, or a renewal to get to it.
The frontier, by the numbers
Every model below runs on Shakudo on 8x AMD MI325X nodes with FP8 quantization and vLLM serving. Numbers are from the vendor-published benchmark suites.
753B params
Set the open-weights bar for agentic coding.
2.8T total / 104B active
Largest open-weights release ever at 2.8T parameters.
2.4T total / 95B active
Alibaba's 2.4T answer to Kimi K3.
320B total / 18B active
DeepSWE 63.4 confirmed on the public leaderboard; beats GLM-5.2 at a tenth of the price.
125B total / 6B active
Hybrid attention design, Qwen4 architecture preview.
743B params
Flagship weights released today after a two-week safety review; the Flash kept the MIT license.
A leading oil and gas producer runs most of its LLM workloads on-prem. “We didn't want our data exposed to laws that we had no control over,” says their Director of Business Intelligence. Frontier models fill the gap where on-prem weight class isn't enough yet, and the data still stays home.
Executive brief
Closed API pricing versus frontier open models on your own GPUs, worked out on your volume. Sent to your inbox.
Thirty minutes on your own data and hardware, with the team that built it.
Book a demo