Enterprise-grade Xeon+GPU AI inference platform — built on the Datafabrix Intel Xeon Server Card, optimised for LLM serving, vision inference, and RAG pipelines.
A 2U inference server combining Intel Xeon compute, DDR5 memory, GPU accelerators, and NVMe local storage — with the manageability and reliability enterprise IT teams need.
Deploy as a standalone inference node or as part of a horizontally-scaled inference cluster. Native integration with the Datafabrix DCIM Platform for fleet-wide visibility.

The block diagram traces the data path of a live inference request through the AI Inference Server. Client requests arrive via the 25/100 GbE front-end network, are load-balanced onto Intel Xeon CPU cores, and passed to the inference runtime (Triton, vLLM, or OpenVINO). The runtime pins model weights to GPU/accelerator memory over the internal PCIe Gen4 fabric, batches requests, and streams token responses back to the client. NVMe storage acts as the model cache and RAG feature store — kept close to the accelerators by the same PCIe fabric, so cold model swap-in is measured in seconds, not minutes.
The connection diagram shows how a DF-S09 chassis physically wires into a customer rack. Two front-end 25/100 GbE ports uplink into the top-of-rack switch for user-facing traffic. A separate 10/25 GbE management port carries out-of-band telemetry to the Datafabrix DCIM stack. Redundant PSUs draw from A+B PDU feeds. Internally, the Xeon compute node connects to GPUs and NVMe pools over the PCIe Gen4 fabric backplane — no external cabling required. Use this diagram to plan rack elevation, cable runs, and network VLAN assignments.
| Interface | Details |
|---|---|
| CPU | Intel Xeon (LGA socket) |
| Memory | 8× DDR5 DIMM slots |
| PCIe expansion | Multiple PCIe Gen5 slots for GPUs/NICs |
| Local storage | M.2 NVMe boot + expansion NVMe |
| Management | ASPEED BMC (Redfish/IPMI over LAN) |
| Networking | Dual RJ45 LAN + optional 10G/25G NICs |
Modern AI inference — LLM serving, real-time computer vision, RAG pipelines — needs a purpose-built platform. High single-thread CPU performance for token pre-processing and post-processing. Enough memory bandwidth to feed the GPUs. Fast local storage for model artefacts. Robust out-of-band management.
The Datafabrix AI Inference Server delivers all of that on our Intel Xeon Server Card. The Xeon CPU handles orchestration and CPU-side operators. Eight DDR5 DIMM slots deliver the memory bandwidth. Multiple PCIe Gen5 slots host GPU accelerators. NVMe local storage keeps models resident.
Integrated ASPEED BMC provides Redfish-standard out-of-band management, making the server first-class in any modern DCIM stack — including our own Datafabrix DCIM Platform.
Deploy internal AI APIs on enterprise-managed hardware.
Cloud/MSP offerings running on standardised inference nodes.
Low-tail-latency inference at the edge of the network.
| Category | Specification |
|---|---|
| Form factor | 2U rack server |
| CPU | Intel Xeon (LGA socket) |
| Memory | 8× DDR5 DIMM slots |
| PCIe | PCIe Gen5 accelerator slots |
| Storage | M.2 NVMe boot + expansion NVMe |
| Networking | Dual RJ45 + optional 10G/25G NICs |
| Management | ASPEED BMC · DCIM-native |
Where it fits in the rack, how it connects, and what you get after installation.
Xeon Server Card + GPU accelerators + high-speed NVMe — turnkey AI inference node. The solution slots into a standard 19-inch rack alongside your existing servers, storage, and network fabric — no re-architecture required.
Target deployments: Enterprise AI teams deploying inference at datacenter scale.
Slide the DF-S09 chassis into a standard 19-inch rack slot. Rail kit included. Cable to your PDU and management network.
Install the Datafabrix PCIe host adapter in your existing server. Connect via PCIe cable to the DF-S09 backplane uplink.
Insert NVMe drives, GPUs, or accelerators into hot-plug slots. Each device trains and appears as a native PCIe endpoint on the host.
Optional: connect to Datafabrix DCIM for fleet-wide visibility. Add more chassis for horizontal scale — no host reconfiguration.
Direct rack integration alongside your existing production servers. Turnkey in a single deployment window.
Scale from a single chassis pilot to full-rack production with zero architectural change.
OEM-ready configuration for organizations building their own branded infrastructure products.
Engineering pilots, reference designs, and OEM co-design programs — we work with your team from concept to production.