DF-S09 · SolutionShipping
XEON-BASED · AI INFERENCE

AI Inference Server.

Enterprise-grade Xeon+GPU AI inference platform — built on the Datafabrix Intel Xeon Server Card, optimised for LLM serving, vision inference, and RAG pipelines.

Ships inside DF-R2
OVERVIEW

Inference infrastructure, built with enterprise discipline.

A 2U inference server combining Intel Xeon compute, DDR5 memory, GPU accelerators, and NVMe local storage — with the manageability and reliability enterprise IT teams need.

Deploy as a standalone inference node or as part of a horizontally-scaled inference cluster. Native integration with the Datafabrix DCIM Platform for fleet-wide visibility.

Built on Intel Xeon Server Card Shipping today
AI Inference Server.
ARCHITECTURE

Solution block diagram.

The block diagram traces the data path of a live inference request through the AI Inference Server. Client requests arrive via the 25/100 GbE front-end network, are load-balanced onto Intel Xeon CPU cores, and passed to the inference runtime (Triton, vLLM, or OpenVINO). The runtime pins model weights to GPU/accelerator memory over the internal PCIe Gen4 fabric, batches requests, and streams token responses back to the client. NVMe storage acts as the model cache and RAG feature store — kept close to the accelerators by the same PCIe fabric, so cold model swap-in is measured in seconds, not minutes.

Xeon CPULGA DDR5 Memory8× DIMMs PCIe Gen5Expansion GPUAccelerator NVMeLocal
DEPLOYMENT TOPOLOGY

Connection diagram.

The connection diagram shows how a DF-S09 chassis physically wires into a customer rack. Two front-end 25/100 GbE ports uplink into the top-of-rack switch for user-facing traffic. A separate 10/25 GbE management port carries out-of-band telemetry to the Datafabrix DCIM stack. Redundant PSUs draw from A+B PDU feeds. Internally, the Xeon compute node connects to GPUs and NVMe pools over the PCIe Gen4 fabric backplane — no external cabling required. Use this diagram to plan rack elevation, cable runs, and network VLAN assignments.

Client Traffic (API) Intel Xeon Server Card Xeon Server Card · Gen5 PCIe GPU 1 GPU 2 NVMe 10G NIC
INTERFACE DETAILS

Interfaces & connectivity.

InterfaceDetails
CPUIntel Xeon (LGA socket)
Memory8× DDR5 DIMM slots
PCIe expansionMultiple PCIe Gen5 slots for GPUs/NICs
Local storageM.2 NVMe boot + expansion NVMe
ManagementASPEED BMC (Redfish/IPMI over LAN)
NetworkingDual RJ45 LAN + optional 10G/25G NICs
HOW IT WORKS

Detailed description.

Modern AI inference — LLM serving, real-time computer vision, RAG pipelines — needs a purpose-built platform. High single-thread CPU performance for token pre-processing and post-processing. Enough memory bandwidth to feed the GPUs. Fast local storage for model artefacts. Robust out-of-band management.

The Datafabrix AI Inference Server delivers all of that on our Intel Xeon Server Card. The Xeon CPU handles orchestration and CPU-side operators. Eight DDR5 DIMM slots deliver the memory bandwidth. Multiple PCIe Gen5 slots host GPU accelerators. NVMe local storage keeps models resident.

Integrated ASPEED BMC provides Redfish-standard out-of-band management, making the server first-class in any modern DCIM stack — including our own Datafabrix DCIM Platform.

Applications

  • LLM inference — Serve large language models with predictable latency.
  • Computer vision — Real-time image and video inference.
  • RAG pipelines — Retrieval + LLM inference in one node.
  • Speech models — STT and TTS at scale.
  • Recommendation systems — Online ranking and personalisation.
USE CASES

Where this solution shines.

Enterprise AI services

Deploy internal AI APIs on enterprise-managed hardware.

AI-as-a-service

Cloud/MSP offerings running on standardised inference nodes.

Latency-sensitive apps

Low-tail-latency inference at the edge of the network.

SPEC HIGHLIGHTS

Technical specifications.

CategorySpecification
Form factor2U rack server
CPUIntel Xeon (LGA socket)
Memory8× DDR5 DIMM slots
PCIePCIe Gen5 accelerator slots
StorageM.2 NVMe boot + expansion NVMe
NetworkingDual RJ45 + optional 10G/25G NICs
ManagementASPEED BMC · DCIM-native
TARGET INDUSTRIES

Who deploys this solution.

Enterprise AICloud AIAI-as-a-ServiceFinancial ServicesRetail
HOW TO DEPLOY · DF-S09

Deploying AI Inference Server in your infrastructure.

Where it fits in the rack, how it connects, and what you get after installation.

AI Inference Server — deployed in a production datacenter row.
Deployment illustration · Ai Inference Server · DF-S09
42U Rack — customer datacenter Existing Host Servers / Compute DF-S09 · AI Inference Server PCIe Gen4 Fabric · Hot-plug ◀── This solution ──▶ Existing Storage Existing Network Ethernet Mgmt · Datafabrix DCIM · Fleet Control
RACK-LEVEL ARCHITECTURE

Where DF-S09 sits in your rack.

Xeon Server Card + GPU accelerators + high-speed NVMe — turnkey AI inference node. The solution slots into a standard 19-inch rack alongside your existing servers, storage, and network fabric — no re-architecture required.

Target deployments: Enterprise AI teams deploying inference at datacenter scale.

  • Cabling: standard PCIe host adapter → PCIe cable → DF-S09 chassis
  • Power: standard C13/C14 rack PDU, 1+1 redundant PSU option
  • Management: 10 GbE + IPMI/BMC out-of-band + Datafabrix DCIM integration
  • Cooling: fits standard rack thermal envelope, no liquid cooling required
01

Install in rack

Slide the DF-S09 chassis into a standard 19-inch rack slot. Rail kit included. Cable to your PDU and management network.

02

Connect PCIe host

Install the Datafabrix PCIe host adapter in your existing server. Connect via PCIe cable to the DF-S09 backplane uplink.

03

Populate the fabric

Insert NVMe drives, GPUs, or accelerators into hot-plug slots. Each device trains and appears as a native PCIe endpoint on the host.

04

Manage & scale

Optional: connect to Datafabrix DCIM for fleet-wide visibility. Add more chassis for horizontal scale — no host reconfiguration.

🎯

Production inference fleets

Direct rack integration alongside your existing production servers. Turnkey in a single deployment window.

📈

AI-as-a-service platforms

Scale from a single chassis pilot to full-rack production with zero architectural change.

🔧

Enterprise LLM hosting

OEM-ready configuration for organizations building their own branded infrastructure products.

Ready to design a rack around it?

Engineering pilots, reference designs, and OEM co-design programs — we work with your team from concept to production.