In this briefing
  1. 01The roadmap proposes a heterogeneous division between GPU prefill and XPU decode
  2. 02Regional residency moves from a contractual requirement into deployment configuration
  3. 03The data boundary needs to extend across compute, storage, networking and governance
  4. 04Long-running tasks belong to a different class of execution duration and environment
  5. What to watch next
  6. Sources and verification
Key points
  1. d-Matrix plans to integrate the Raptor XPU into NVIDIA MGX and NVLink Fusion, using GPUs for prefill and an XPU to accelerate decode as an example of heterogeneous inference. Raptor has not yet taped out, and initial availability of the integrated rack is only expected in the fourth quarter of 2027, so it cannot be treated as delivered performance.
  2. Baseten allows every replica in a deployment to be restricted to a specified Baseten region; NVIDIA, Palantir and Dell set out deployment forms for the same class of enterprise workflow across cloud, on-premises and isolated environments.
  3. The OpenAI Agents API places sessions, context and the execution loop for long-running tasks in a managed service, while allowing a choice of hosted, VPC or partner sandboxes; the product remains in public beta.
Signal 01

The roadmap proposes a heterogeneous division between GPU prefill and XPU decode

At 20:59:26 on 10 September (Beijing time), d-Matrix published a multi-year roadmap with NVIDIA; NVIDIA then published its corresponding account at 21:00:21. The two materials place the Raptor XPU in an NVIDIA MGX rack: NVLink handles scale-up within the rack, Spectrum-X handles scale-out, and the plan also connects Vera CPU, ConnectX-9 SuperNIC and BlueField-4 DPU. NVIDIA stresses that the common rack also encompasses power, cooling and the supply chain, rather than providing only a chip interconnect.

d-Matrix lists AI coding, real-time chat and voice Agents as target low-latency scenarios, and gives an example of heterogeneous disaggregation: NVIDIA GPU hardware handles compute-intensive prefill, while the Raptor XPU accelerates interaction-sensitive decode. Raptor uses a 3D-stacking approach combining a DRAM memory chip with an SRAM compute chip, but the announcement gives no figures for memory capacity, bandwidth, power consumption, first-token latency, per-token latency, throughput or concurrency. The previous-generation Corsair is the product currently in production; Raptor is expected to reach tape-out by the end of 2026, with initial availability of its MGX integration expected in the fourth quarter of 2027.

What this may mean for enterprise adoption

This roadmap shows why ‘low latency’ cannot be reduced to a single average response time. Capacity planning and procurement testing therefore cover prefill, decode, batch throughput, tail latency and inter-stage transfer overhead, and must distinguish roadmap hardware from hardware that is already available. Vendor descriptions of ultralow latency, energy efficiency and lower risk are not accompanied by reproducible benchmarks; at this stage they can only be treated as an architectural direction, not counted in production capacity ahead of delivery.

Signal 02

Regional residency moves from a contractual requirement into deployment configuration

At 02:56:54 on 11 September (Beijing time), Baseten released Regional Deployments, allowing every replica in a deployment to be restricted to a particular Baseten region. Users can select the region through the dashboard, CLI or Management API; the official example is `baseten model push --region eu`. The announcement lists compliance, data residency and lower latency as reasons to use the capability.

This update establishes that replicas can be bound to regions defined by Baseten, but it does not list every available region, availability-zone granularity, cross-region failover, network paths, pricing or regional service levels. Nor does it say whether control-plane logs, backups and support data follow exactly the same geographical boundary as inference replicas.

What this may mean for enterprise adoption

For regulated data or cross-region services, this makes region a distinct field in a versioned deployment inventory, linked to the model revision, replica count, disaster-recovery policy and permitted data categories. Selecting a region can shorten network distance and help meet residency requirements, but ‘the workload runs in a region’ and ‘related data never leaves that jurisdiction’ remain separate judgements. End-to-end compliance also depends on failover, logs, backups, support access and contractual boundaries.

Signal 03

The data boundary needs to extend across compute, storage, networking and governance

At 17:00 on 10 September (Beijing time), NVIDIA and Palantir announced a supply-chain AI collaboration: customised Nemotron open models enter Palantir Foundry and AIP, with business objects and relationships represented through Palantir Ontology. Customers can post-train models with their own data; cuOpt is used for constraint and scenario analysis, while NeMo AutoModel, NeMo RL and Palantir Autopilot are used to incorporate business feedback into continuous improvement. The announcement explicitly retains expert control over final decisions, says that the stack can run in the cloud or on-premises, and states that it will first be deployed in NVIDIA's own supply chain.

At 21:38:53 (Beijing time), Dell added an on-premises reference form: PowerEdge XE9780 with Blackwell Ultra GPU accelerators handles model customisation, retrieval, inference and Agent workloads, while PowerEdge R670 handles the control plane and data processing. ObjectScale, PowerFlex, Spectrum-X, ConnectX and PowerSwitch cover object storage, block storage, GPU networking and management-traffic isolation respectively. Dell also includes Rubix multi-tenant policies and Apollo deployment and lifecycle management across enterprise data centres, sovereign facilities and air-gapped environments in the architecture, but discloses no overall pricing, delivery lead time, power consumption or business outcomes for the solution.

What this may mean for enterprise adoption

Regional residency is only one layer of a data boundary. An on-premises deployment must also answer which forms of storage hold models and data, how GPU and management traffic are isolated, who can update the runtime, how feedback data enters post-training, and who is responsible when a failure occurs. NVIDIA's and Dell's materials are vendor solutions and forward-looking statements: they can define components to be accepted, but cannot prove that every industry customer has already obtained the same outcomes or complete air-gap capability.

Signal 04

Long-running tasks belong to a different class of execution duration and environment

At 04:54:46 on 11 September (Beijing time), a staff announcement on the official OpenAI Developer Community said that the Agents API had entered public beta for all developers. OpenAI manages the Agent loop, model calls, tool use, long-running sessions and context management, while developers still decide the Agent's capabilities and where code and files run. Environment options include a self-provided sandbox, partner sandboxes, an OpenAI-hosted sandbox, and managed or VPC deployments with options for CPU, GPU, memory, files and secret storage.

The announcement says that the Agents API itself carries no additional charge, with costs arising from the tokens and tools used; this is not a complete task cost, performance guarantee or service level. It also gives no maximum execution time, recovery-point semantics, concurrency limit or failure-retry guarantee. The product article is dated only 10 September, so this article uses the precise publication time of the official forum announcement. That time establishes that the announcement falls within the observation window, but cannot be used to infer the product article's exact publication time.

What this may mean for enterprise adoption

Long-running Agents and real-time conversations have different operational objectives: the former place greater emphasis on state persistence, permissions, resource limits, recovery and artefact boundaries, while the latter place greater emphasis on first-token and per-token latency. When both enter the same service catalogue, their acceptance metrics still need to remain separate. Especially during public beta, recoverable state, timeout rules and spending limits for long-running tasks remain operational boundaries that are not replaced by a product announcement.

Verification

Sources and verification

  1. d-Matrix Adopts NVIDIA NVLink Fusion for Rack-Scale XPU DeploymentNVIDIA · 2026-09-10T13:00:21Z · Official announcement
  2. Regional deploymentsBaseten · 2026-09-10T19:56:54+01:00 · Official announcement
  3. NVIDIA and Palantir Bring Sovereign Intelligence to Critical Supply ChainsNVIDIA / Palantir · 2026-09-10T09:00:00Z · Official announcement
  4. Sovereign AI for the Supply Chain, Built on Infrastructure You ControlDell Technologies · 2026-09-10T13:38:53+00:00 · Official announcement
  5. Introducing the Agents API and hosted sandboxesOpenAI Developer Community · 2026-09-10T20:54:46.591Z · Official announcement

Golden Data has edited this briefing from the public materials listed above. The original sources govern facts and figures. The enterprise relevance sections are Golden Data editorial analysis and do not constitute an endorsement of any third-party product.

← Back to AI Daily Briefing