In this briefing
- 01The roadmap proposes a heterogeneous division between GPU prefill and XPU decode
- 02Regional residency moves from a contractual requirement into deployment configuration
- 03The data boundary needs to extend across compute, storage, networking and governance
- 04Long-running tasks belong to a different class of execution duration and environment
- →What to watch next
- ↗Sources and verification
- d-Matrix plans to integrate the Raptor XPU into NVIDIA MGX and NVLink Fusion, using GPUs for prefill and an XPU to accelerate decode as an example of heterogeneous inference. Raptor has not yet taped out, and initial availability of the integrated rack is only expected in the fourth quarter of 2027, so it cannot be treated as delivered performance.
- Baseten allows every replica in a deployment to be restricted to a specified Baseten region; NVIDIA, Palantir and Dell set out deployment forms for the same class of enterprise workflow across cloud, on-premises and isolated environments.
- The OpenAI Agents API places sessions, context and the execution loop for long-running tasks in a managed service, while allowing a choice of hosted, VPC or partner sandboxes; the product remains in public beta.
The roadmap proposes a heterogeneous division between GPU prefill and XPU decode
At 20:59:26 on 10 September (Beijing time), d-Matrix published a multi-year roadmap with NVIDIA; NVIDIA then published its corresponding account at 21:00:21. The two materials place the Raptor XPU in an NVIDIA MGX rack: NVLink handles scale-up within the rack, Spectrum-X handles scale-out, and the plan also connects Vera CPU, ConnectX-9 SuperNIC and BlueField-4 DPU. NVIDIA stresses that the common rack also encompasses power, cooling and the supply chain, rather than providing only a chip interconnect.
d-Matrix lists AI coding, real-time chat and voice Agents as target low-latency scenarios, and gives an example of heterogeneous disaggregation: NVIDIA GPU hardware handles compute-intensive prefill, while the Raptor XPU accelerates interaction-sensitive decode. Raptor uses a 3D-stacking approach combining a DRAM memory chip with an SRAM compute chip, but the announcement gives no figures for memory capacity, bandwidth, power consumption, first-token latency, per-token latency, throughput or concurrency. The previous-generation Corsair is the product currently in production; Raptor is expected to reach tape-out by the end of 2026, with initial availability of its MGX integration expected in the fourth quarter of 2027.
This roadmap shows why ‘low latency’ cannot be reduced to a single average response time. Capacity planning and procurement testing therefore cover prefill, decode, batch throughput, tail latency and inter-stage transfer overhead, and must distinguish roadmap hardware from hardware that is already available. Vendor descriptions of ultralow latency, energy efficiency and lower risk are not accompanied by reproducible benchmarks; at this stage they can only be treated as an architectural direction, not counted in production capacity ahead of delivery.
Regional residency moves from a contractual requirement into deployment configuration
At 02:56:54 on 11 September (Beijing time), Baseten released Regional Deployments, allowing every replica in a deployment to be restricted to a particular Baseten region. Users can select the region through the dashboard, CLI or Management API; the official example is `baseten model push --region eu`. The announcement lists compliance, data residency and lower latency as reasons to use the capability.
This update establishes that replicas can be bound to regions defined by Baseten, but it does not list every available region, availability-zone granularity, cross-region failover, network paths, pricing or regional service levels. Nor does it say whether control-plane logs, backups and support data follow exactly the same geographical boundary as inference replicas.
For regulated data or cross-region services, this makes region a distinct field in a versioned deployment inventory, linked to the model revision, replica count, disaster-recovery policy and permitted data categories. Selecting a region can shorten network distance and help meet residency requirements, but ‘the workload runs in a region’ and ‘related data never leaves that jurisdiction’ remain separate judgements. End-to-end compliance also depends on failover, logs, backups, support access and contractual boundaries.
The data boundary needs to extend across compute, storage, networking and governance
At 17:00 on 10 September (Beijing time), NVIDIA and Palantir announced a supply-chain AI collaboration: customised Nemotron open models enter Palantir Foundry and AIP, with business objects and relationships represented through Palantir Ontology. Customers can post-train models with their own data; cuOpt is used for constraint and scenario analysis, while NeMo AutoModel, NeMo RL and Palantir Autopilot are used to incorporate business feedback into continuous improvement. The announcement explicitly retains expert control over final decisions, says that the stack can run in the cloud or on-premises, and states that it will first be deployed in NVIDIA's own supply chain.
At 21:38:53 (Beijing time), Dell added an on-premises reference form: PowerEdge XE9780 with Blackwell Ultra GPU accelerators handles model customisation, retrieval, inference and Agent workloads, while PowerEdge R670 handles the control plane and data processing. ObjectScale, PowerFlex, Spectrum-X, ConnectX and PowerSwitch cover object storage, block storage, GPU networking and management-traffic isolation respectively. Dell also includes Rubix multi-tenant policies and Apollo deployment and lifecycle management across enterprise data centres, sovereign facilities and air-gapped environments in the architecture, but discloses no overall pricing, delivery lead time, power consumption or business outcomes for the solution.
Regional residency is only one layer of a data boundary. An on-premises deployment must also answer which forms of storage hold models and data, how GPU and management traffic are isolated, who can update the runtime, how feedback data enters post-training, and who is responsible when a failure occurs. NVIDIA's and Dell's materials are vendor solutions and forward-looking statements: they can define components to be accepted, but cannot prove that every industry customer has already obtained the same outcomes or complete air-gap capability.
Long-running tasks belong to a different class of execution duration and environment
At 04:54:46 on 11 September (Beijing time), a staff announcement on the official OpenAI Developer Community said that the Agents API had entered public beta for all developers. OpenAI manages the Agent loop, model calls, tool use, long-running sessions and context management, while developers still decide the Agent's capabilities and where code and files run. Environment options include a self-provided sandbox, partner sandboxes, an OpenAI-hosted sandbox, and managed or VPC deployments with options for CPU, GPU, memory, files and secret storage.
The announcement says that the Agents API itself carries no additional charge, with costs arising from the tokens and tools used; this is not a complete task cost, performance guarantee or service level. It also gives no maximum execution time, recovery-point semantics, concurrency limit or failure-retry guarantee. The product article is dated only 10 September, so this article uses the precise publication time of the official forum announcement. That time establishes that the announcement falls within the observation window, but cannot be used to infer the product article's exact publication time.
Long-running Agents and real-time conversations have different operational objectives: the former place greater emphasis on state persistence, permissions, resource limits, recovery and artefact boundaries, while the latter place greater emphasis on first-token and per-token latency. When both enter the same service catalogue, their acceptance metrics still need to remain separate. Especially during public beta, recoverable state, timeout rules and spending limits for long-running tasks remain operational boundaries that are not replaced by a product announcement.
What to watch next
- Whether d-Matrix publishes like-for-like benchmarks for first-token latency, per-token latency, throughput, power consumption and inter-stage transfer after Raptor tape-out and MGX integration.
- Whether Baseten publishes the regional deployment region list, failover behaviour, log and backup residency scope, and pricing and service levels for each region.
- Whether enterprises can use one runtime inventory to record latency class, region, on-premises or cloud form, data category, execution duration, recovery semantics and spending limits together.
Sources and verification
Golden Data has edited this briefing from the public materials listed above. The original sources govern facts and figures. The enterprise relevance sections are Golden Data editorial analysis and do not constitute an endorsement of any third-party product.
← Back to AI Daily Briefing