DF-R2Shipping

Inference Rack.

Serve production LLM and vision-model traffic at enterprise scale. Cold-swap models in seconds. Keep the whole fleet observable from the same console.

AT A GLANCE

Inference Rack — what you get.

A single Datafabrix DF-R2 rack, delivered as one purchase order, installed as one unit, supported as one system.

Mission
Production LLM serving at enterprise scale
Form factor
42U · full rack · 18 kW nominal · N+1 PSU
Workload profile
Production LLM serving · RAG inference · Vision inference · Multi-tenant model hosting · Enterprise chatbot backends
Target customer
Enterprises deploying inference at scale, AI platform teams, industry-cloud providers, regulated verticals (finance, healthcare, government)
KPIS

The numbers that matter.

The specific metrics our customers ask about first — and where this rack is engineered to win.

Model cold-swap time
Sub-second from NVMe cache
Concurrent models hosted
Dozens per rack, tenancy-aware
Token throughput headroom
Sized to node accelerator count
P99 request latency
Bounded by accelerator, not by fabric
BILL OF MATERIALS

Every layer of the rack, itemised.

Merchant silicon is called out explicitly. Datafabrix engineering is called out explicitly. No hand-wave categories.

Layer Component Notes
Compute 6 × AI Inference Servers (DF-S09) · Intel Xeon + accelerators Xeon control plane · GPU or dedicated inference accelerator per node
Fabric 1 × Datafabrix PCIe Gen4 Fabric Backplane (DF-B4) Fabric-attached model cache and RAG feature store
Model cache AI Storage Appliance (DF-S05) · pre-validated Model weights and RAG indexes on NVMe — sub-second cold-swap
Controller Storage Controller Node (DF-S12) NVMe-oF and S3 endpoints for client-facing traffic
DCIM All 8 DCIM modules pre-instrumented Includes StorageOS for model-cache SSD wear tracking
Networking Front-end 2 × 100 GbE + management 25 GbE HA pair leaf uplinks; dedicated OOB management
Cooling Air-cooled · standard datacenter thermal envelope Fits any Tier III colo. No liquid required.
WHY THIS RACK

The problems this rack addresses.

Inference is not just small training

Production inference has completely different constraints — latency SLAs, multi-tenancy, model-swap velocity, and cost-per-token. This rack was designed for those metrics, not adapted from a training design.

Model cold-swap is where enterprises stall

Every serious inference platform needs to swap between dozens of models. Doing that from network storage adds seconds; doing it from a fabric-attached NVMe cache is sub-second. That is the difference between a fleet that scales and one that queues.

Multi-tenancy needs observability

When one tenant is degrading another tenant, you need to see it in the console within seconds. Our Vision and Insight modules were built for exactly this — request-level attribution across the fleet.

Weeks matter, not quarters

AI product teams cannot wait eight months for infrastructure — the model they wanted to serve is already three versions old. This rack ships and powers on inside a quarter.

WHERE IT'S DEPLOYED

Four workloads. One rack.

Production LLM serving

Serve chat, completions, and embeddings at enterprise volume with predictable P99.

RAG inference

Feature stores kept on the fabric-attached NVMe tier — sub-millisecond retrieval.

Vision inference

Classification, detection, OCR at millions of requests per hour per rack.

Enterprise chatbot back-end

Multi-tenant, RBAC-aware model hosting for internal AI products.

FROM ORDER TO POWER-ON

From order to power-on. In weeks.

Our delivery discipline is the single biggest reason customers pick us over DIY assembly or a locked AI system. Here is what the timeline actually looks like.

01

Discovery & quote

One call to align on workload, site, and timeline. Written quote with BOM, power/cooling profile, and delivery date inside one business week.

02

Build & validate

Rack assembled and validated at our integration facility. Firmware loaded. DCIM modules provisioned. Burn-in tested against the specific workload class.

03

Ship & install

Rack ships fully assembled. On-site installation by our team or a certified partner. Cabled, powered, network-integrated, DCIM streaming to your operators.

04

Power-on & handover

Workload deployed on the rack. Acceptance tests. Handover to your operations team, with our engineering on standby for the first 90 days.

WHAT'S INSIDE THIS RACK

The solutions inside this rack.

Each Datafabrix rack is a curated combination of solutions. Below are the solutions that ship inside this rack. Click any card for the block diagram, connection topology, and per-subsystem specs.

READY FOR YOUR FIRST DF-R2

Let's design your first inference rack.

Tell us your workload, your site, and your timeline. We will respond within one business day.