Book a demo

The open frontier, on your own hardware

Frontier-class AI inside your company

The newest open models now match the big closed vendors on coding work, at a fraction of the cost. Shakudo runs them on your hardware, and your data never leaves your network. See the full roster, benchmarked side by side below.

By submitting your email you agree to receive occasional updates from Shakudo. See our Privacy Policy.

  • No data leaves the network
  • No per-token bill
  • Fine-tune to your data

Frontier-class results, on your own hardware.

The open models at the frontier now hold their own against the closed ones on the benchmarks your team actually cares about. You get that capability without the API, the vendor, or the exposure.

The token bill drops off the top.

Open-weight models are priced in compute, not per token. Run them on hardware you already own and the per-1,000-token rate stops climbing with every request you make.

Your data stays yours.

Prompts and outputs never leave the network. No training on your data, no retention on their side, no foreign legal framework you have no control over.

Fine-tune to your own data.

Open weights mean you can steer the model. Point it at your documents, your tone, your domain, and it gets better at your work instead of guessing at everyone else's.

One API, any model.

Shakudo fronts the whole roster behind a single OpenAI-compatible endpoint. Swap models as the frontier moves, and route each task to the one that does it best.

The frontier keeps moving.

Closed-model capability plateaus and then prices itself up. The open line is still rising, and it does not require a seat, a contract, or a renewal to get to it.

The frontier, by the numbers

Six months of open-weights releases, benchmarked side-by-side.

Every model below runs on Shakudo on 8x AMD MI325X nodes with FP8 quantization and vLLM serving. Numbers are from the vendor-published benchmark suites.

June 2026

GLM-5.2

Z.ai

753B params

  • SWE-bench Pro 62.1
  • AIME 2026 99.2
  • Terminal-Bench 2.1 81.0

Set the open-weights bar for agentic coding.

July 2026

Kimi-K3

Moonshot AI

2.8T total / 104B active

  • Class First open 3T
  • Quant MXFP4

Largest open-weights release ever at 2.8T parameters.

July 2026

Qwen3.8-Max

Alibaba

2.4T total / 95B active

  • SWE-bench Pro 67.7
  • Terminal-Bench 86.6

Alibaba's 2.4T answer to Kimi K3.

August 2026

GLM-5.3-Flash

Z.ai

320B total / 18B active

  • DeepSWE 63.4
  • Terminal-Bench 88.2
  • Toolathlon 73.0

DeepSWE 63.4 confirmed on the public leaderboard; beats GLM-5.2 at a tenth of the price.

August 2026

Qwen3.8-Flash-Next

Alibaba

125B total / 6B active

  • LiveCodeBench v6 91.9
  • GPQA Diamond 91.7

Hybrid attention design, Qwen4 architecture preview.

August 2026

GLM-5.3

Z.ai

743B params

  • License Semi-non-commercial
  • FP8 weights ~756 GB

Flagship weights released today after a two-week safety review; the Flash kept the MIT license.

A leading oil and gas producer runs most of its LLM workloads on-prem. “We didn't want our data exposed to laws that we had no control over,” says their Director of Business Intelligence. Frontier models fill the gap where on-prem weight class isn't enough yet, and the data still stays home.

ON-PREM
For the majority of LLM workloads
1 TB/MO
Of data governed on the stack
375K
Boe/d the analytics support
<1 HR
For a business user to build new analytics
Read the case study ›

Executive brief

See what the token bill looks like at your scale.

Closed API pricing versus frontier open models on your own GPUs, worked out on your volume. Sent to your inbox.

See it on your workloads.

Thirty minutes on your own data and hardware, with the team that built it.

Book a demo