DEPLOYMENT ASSURANCE / EXPERT-ROUTED AI

Prove it before
you deploy.

Plan an exact deployment. Run the open gates. Verify the resulting evidence. Share a portable receipt whose limits remain attached.

LOCAL DEMO Deterministic fixture bytes · no upload · no model-performance claim

15 sourced claims0 admitted controlled runs0 fabricated scores

Independent deployment intelligence from LockedIn Labs

PINNED ARTIFACT MAP REGISTRY FACTS / V1.0.0
RROUTER
E01
E02
E03
E04
E05
E06
E07
E08
8 OF 128 EXPERTS ACTIVE / TOKEN
8 × H200 NODE VIEW1 GPU MIN. / CHECKPOINT ONLY
PINNED ARTIFACT61.1 GB
EXPERT ROUTING8 / 128
STATIC LOWER BOUND1 GPU MIN.

Qwen3-30B-A3B · sourced registry facts + calculated static fit · topology illustration, not measured runtime data

PLAN · EXACT ARTIFACTRUN · DEPLOYBENCH V0.1VERIFY · PASSPORT V0.2SHARE · PORTABLE RECEIPT

ONE SYSTEM / MULTIPLE ENTRY POINTS

PLATFORM V0.1

Enter wherever the decision starts.

The website is the visual workbench. The protocol, runner and developer surfaces make the same decision system portable.

THE DEPLOYMENT GAP

01 / 05

37B active does not mean 37B resident.

Sparse activation reduces compute. It does not erase model weights, KV cache, runtime buffers, replication, or all-to-all traffic.

MOEModels separates can load, can run, can scale, and makes economic sense— because they are four different answers.

01
FACTS

Every model specification links back to primary evidence.

02
CALCULATIONS

Every derived value exposes its exact inputs and formula.

03
DECISIONS

Every result ends with a bottleneck and a next action.

MODEL INTELLIGENCE

02 / 05

Five artifact-exact
MoE records. One evidence layer.

Sourced architecture and hardware facts, with licensing, runtime compatibility and deployment unknowns kept visible.

Open model index
01 Primary source

Moonshot AI

Kimi K3

A native multimodal 2.8T-class model for long-horizon coding, knowledge work, and million-token contexts.

TOTAL2.8T
ACTIVE104B
EXPERTS896
Long-horizon agentsView intelligence
02 Primary source

DeepSeek

DeepSeek V4 Pro

A frontier-scale model combining fine-grained sparse experts with a million-token context for agentic work.

TOTAL1.6T
ACTIVE49B
EXPERTS384
Frontier open reasoningView intelligence
03 Primary source

Z.ai

GLM-5.2

An open flagship designed for long engineering trajectories across a sourced million-token context.

TOTAL753B
ACTIVEUnknown
EXPERTS256
Long-horizon codingView intelligence
04 Primary source

Google

Gemma 4 26B A4B IT

A compact multimodal MoE with 25.2B total and 3.8B active parameters, designed for accessible inference.

TOTAL25.2B
ACTIVE3.8B
EXPERTS128
Single-accelerator baselineView intelligence

Specifications shown above are sourced from pinned manifests, configurations, model cards, or technical releases. No benchmark ranking is implied.

MOE FIT CHECK

03 / 05

Turn an exact artifact into an honest lower bound.

Select a pinned checkpoint, hardware target, topology, and declared reserve. See what is proven—and what remains unknown.

Exact manifest bytes + integer math. No generated throughput, capacity, or price.

Deterministic registry engine

Will the checkpoint fit?

Test an exact, pinned artifact against advertised accelerator memory—then keep every unsupported conclusion visibly unknown.

Sourced + calculated

01 · Define baseline

Residency inputs

5 inputs
Total2.8T
Active / token104B
Artifact1.6 TB
GPUs per node

02 · Fit decision

Checkpoint cannot fit

Kimi K3 · 896 / top-16 experts · 1M context · 8 × NVIDIA H200 SXM · 13% reserve

Decision boundaryCheckpoint bytes only
01Baseline fails

Load

The checkpoint alone requires at least 13 GPUs at this reserve, before runtime allocations.

02Evidence required

Run

A baseline pass does not prove loader, kernel, quantization, sharding, or expert-parallel support.

03Unknown

Scale

No measured workload profile is attached, so throughput, latency, KV demand, and skew are not projected.

04Unknown

Economics

No dated provider, region, or utilization record is selected, so the engine emits no invented cost.

03 · Explainable topology

Requested

Pure integer math
Artifact
1.6 TBtensors
Reserve
Experts
Compute8 × H200
Checkpoint tensors
1.6 TB
exact manifest metadata
Usable / GPU
122.7 GB
after declared reserve
Device lower bound
13
before node rounding
Node-rounded
16 GPUs
2 × 8-GPU nodes

RequestedThe topology you asked the engine to test against the checkpoint baseline.

View calculation anatomy
Exact artifact tensor bytes
1,560,860,324,864
Advertised memory / GPU
141 GB
Declared reserve / GPU
18.3 GB
Usable memory / GPU
122.7 GB

usable = floor(advertised bytes × (10,000 − reserve bps) / 10,000); minimum GPUs = ceil(checkpoint tensor bytes / usable bytes). A failure is conclusive for this no-offload baseline; a pass is only a candidate for runtime validation.

Artifact evidence: Kimi K3 pinned tensor manifest

SOURCED + CALCULATED · REGISTRY V1.0.0

Checkpoint residency is not full runtime residency. Validate loader support, memory use, topology, quality, latency, throughput, and cost on the intended system.

BENCHMARK EVIDENCE

04 / 05

A leaderboard should show its work.

Fifteen owner-reported claims keep model scope, settings, sources and missing context attached. No controlled run has crossed the comparison gate yet.

Enter the evidence lab
EVIDENCE CLASSMETHODSTATUS
Owner reports15 sourced claimsAVAILABLE
Controlled endpoint runs0 admitted packsOPEN
Verified Deployment Passports0 admitted packsOPEN
Independent reproductions0 reproduced packsFUTURE
Calibrated configuration corpus0 configurationsFUTURE
DeployBench measures. Passport verifies the receipt.Admission and reproduction remain separate.

OPEN INFRASTRUCTURE

05 / 05

One engine.
Every interface.

Workbench, CLI, runner, REST API, TypeScript SDK and MCP server share the registry, evidence classes and deterministic planning contract.

moemodels / fitREGISTRY V1
$ npm run moemodels -- fit \
moonshotai/Kimi-K3 \
nvidia/h200-sxm-141gb \
--gpus 8 --reserve-pct 13
CHECKPOINTFAIL × 8
TOPOLOGY13 → 16
RUNTIMEUNKNOWN

Real offline CLI · same registry and integer engine as the hosted Fit Check

RESEARCH DESK

Read past the parameter count.

View all research

FIELD GUIDE 01

Active parameters are not a deployment plan

A practical guide to resident weights, KV cache, runtime overhead, and the four different meanings of “fits.”

11 min read · Methodology

SYSTEMS NOTE 07

Expert parallelism, without the hand-waving

How topology, all-to-all traffic, load imbalance, and hot experts reshape the economics of sparse inference.

14 min read · Infrastructure

BUYER BRIEF 03

Cloud endpoint or private cluster?

A decision framework for comparing utilization, privacy, operational burden, and three-year total cost.

9 min read · Economics

START WITH THE PROOF

Break the receipt.
Watch trust fail closed.

Verify a deterministic Passport fixture, change exactly one byte, and see the content address reject it—all locally in your browser.

Run the 90-second check