Moonshot AI
Kimi K3
A native multimodal 2.8T-class model for long-horizon coding, knowledge work, and million-token contexts.
DEPLOYMENT ASSURANCE / EXPERT-ROUTED AI
Plan an exact deployment. Run the open gates. Verify the resulting evidence. Share a portable receipt whose limits remain attached.
LOCAL DEMO Deterministic fixture bytes · no upload · no model-performance claim
Independent deployment intelligence from LockedIn Labs
Qwen3-30B-A3B · sourced registry facts + calculated static fit · topology illustration, not measured runtime data
ONE SYSTEM / MULTIPLE ENTRY POINTS
PLATFORM V0.1
The website is the visual workbench. The protocol, runner and developer surfaces make the same decision system portable.
Bind an exact artifact to hardware, runtime, workload and SLA. Calculate the static floor and open validation gates.
Measure endpoint TTFT, request latency, success and aggregate throughput without retaining prompts or responses.
Pack compatible trials under one digest, verify summaries and signatures locally, and keep every reproducibility gap attached.
Inspect artifact-exact architecture and owner-reported evidence with every missing field kept visible.
Use the same decision contract through the CLI, benchmark runner, REST API, TypeScript SDK and MCP server.
THE DEPLOYMENT GAP
01 / 05
Sparse activation reduces compute. It does not erase model weights, KV cache, runtime buffers, replication, or all-to-all traffic.
MOEModels separates can load, can run, can scale, and makes economic sense— because they are four different answers.
Every model specification links back to primary evidence.
Every derived value exposes its exact inputs and formula.
Every result ends with a bottleneck and a next action.
MODEL INTELLIGENCE
02 / 05
Sourced architecture and hardware facts, with licensing, runtime compatibility and deployment unknowns kept visible.
Open model indexMoonshot AI
A native multimodal 2.8T-class model for long-horizon coding, knowledge work, and million-token contexts.
DeepSeek
A frontier-scale model combining fine-grained sparse experts with a million-token context for agentic work.
Z.ai
An open flagship designed for long engineering trajectories across a sourced million-token context.
A compact multimodal MoE with 25.2B total and 3.8B active parameters, designed for accessible inference.
Specifications shown above are sourced from pinned manifests, configurations, model cards, or technical releases. No benchmark ranking is implied.
MOE FIT CHECK
03 / 05
Select a pinned checkpoint, hardware target, topology, and declared reserve. See what is proven—and what remains unknown.
Exact manifest bytes + integer math. No generated throughput, capacity, or price.
Deterministic registry engine
Test an exact, pinned artifact against advertised accelerator memory—then keep every unsupported conclusion visibly unknown.
02 · Fit decision
Kimi K3 · 896 / top-16 experts · 1M context · 8 × NVIDIA H200 SXM · 13% reserve
The checkpoint alone requires at least 13 GPUs at this reserve, before runtime allocations.
A baseline pass does not prove loader, kernel, quantization, sharding, or expert-parallel support.
No measured workload profile is attached, so throughput, latency, KV demand, and skew are not projected.
No dated provider, region, or utilization record is selected, so the engine emits no invented cost.
03 · Explainable topology
RequestedThe topology you asked the engine to test against the checkpoint baseline.
usable = floor(advertised bytes × (10,000 − reserve bps) / 10,000); minimum GPUs = ceil(checkpoint tensor bytes / usable bytes). A failure is conclusive for this no-offload baseline; a pass is only a candidate for runtime validation.
Artifact evidence: Kimi K3 pinned tensor manifest ↗
BENCHMARK EVIDENCE
04 / 05
Fifteen owner-reported claims keep model scope, settings, sources and missing context attached. No controlled run has crossed the comparison gate yet.
Enter the evidence labOPEN INFRASTRUCTURE
05 / 05
Workbench, CLI, runner, REST API, TypeScript SDK and MCP server share the registry, evidence classes and deterministic planning contract.
$ npm run moemodels -- fit \
moonshotai/Kimi-K3 \
nvidia/h200-sxm-141gb \
--gpus 8 --reserve-pct 13
Real offline CLI · same registry and integer engine as the hosted Fit Check
RESEARCH DESK
FIELD GUIDE 01
A practical guide to resident weights, KV cache, runtime overhead, and the four different meanings of “fits.”
SYSTEMS NOTE 07
How topology, all-to-all traffic, load imbalance, and hot experts reshape the economics of sparse inference.
BUYER BRIEF 03
A decision framework for comparing utilization, privacy, operational burden, and three-year total cost.
START WITH THE PROOF
Verify a deterministic Passport fixture, change exactly one byte, and see the content address reject it—all locally in your browser.
Run the 90-second check