Serve production LLM and vision-model traffic at enterprise scale. Cold-swap models in seconds. Keep the whole fleet observable from the same console.
A single Datafabrix DF-R2 rack, delivered as one purchase order, installed as one unit, supported as one system.
The specific metrics our customers ask about first — and where this rack is engineered to win.
Merchant silicon is called out explicitly. Datafabrix engineering is called out explicitly. No hand-wave categories.
| Layer | Component | Notes |
|---|---|---|
| Compute | 6 × AI Inference Servers (DF-S09) · Intel Xeon + accelerators | Xeon control plane · GPU or dedicated inference accelerator per node |
| Fabric | 1 × Datafabrix PCIe Gen4 Fabric Backplane (DF-B4) | Fabric-attached model cache and RAG feature store |
| Model cache | AI Storage Appliance (DF-S05) · pre-validated | Model weights and RAG indexes on NVMe — sub-second cold-swap |
| Controller | Storage Controller Node (DF-S12) | NVMe-oF and S3 endpoints for client-facing traffic |
| DCIM | All 8 DCIM modules pre-instrumented | Includes StorageOS for model-cache SSD wear tracking |
| Networking | Front-end 2 × 100 GbE + management 25 GbE | HA pair leaf uplinks; dedicated OOB management |
| Cooling | Air-cooled · standard datacenter thermal envelope | Fits any Tier III colo. No liquid required. |
Production inference has completely different constraints — latency SLAs, multi-tenancy, model-swap velocity, and cost-per-token. This rack was designed for those metrics, not adapted from a training design.
Every serious inference platform needs to swap between dozens of models. Doing that from network storage adds seconds; doing it from a fabric-attached NVMe cache is sub-second. That is the difference between a fleet that scales and one that queues.
When one tenant is degrading another tenant, you need to see it in the console within seconds. Our Vision and Insight modules were built for exactly this — request-level attribution across the fleet.
AI product teams cannot wait eight months for infrastructure — the model they wanted to serve is already three versions old. This rack ships and powers on inside a quarter.
Serve chat, completions, and embeddings at enterprise volume with predictable P99.
Feature stores kept on the fabric-attached NVMe tier — sub-millisecond retrieval.
Classification, detection, OCR at millions of requests per hour per rack.
Multi-tenant, RBAC-aware model hosting for internal AI products.
Our delivery discipline is the single biggest reason customers pick us over DIY assembly or a locked AI system. Here is what the timeline actually looks like.
One call to align on workload, site, and timeline. Written quote with BOM, power/cooling profile, and delivery date inside one business week.
Rack assembled and validated at our integration facility. Firmware loaded. DCIM modules provisioned. Burn-in tested against the specific workload class.
Rack ships fully assembled. On-site installation by our team or a certified partner. Cabled, powered, network-integrated, DCIM streaming to your operators.
Workload deployed on the rack. Acceptance tests. Handover to your operations team, with our engineering on standby for the first 90 days.
Each Datafabrix rack is a curated combination of solutions. Below are the solutions that ship inside this rack. Click any card for the block diagram, connection topology, and per-subsystem specs.
Tell us your workload, your site, and your timeline. We will respond within one business day.