In this briefing
  1. 01Data-centre and edge tier: one architecture family spans a 350W inference card and an edge SoC
  2. 02Appliance tier: the same 150W figure does not make two product states equivalent
  3. 03Desktop prototype tier: AI Cube demonstrates a local combination of 120B and 3B models
  4. What to watch next
  5. Sources and verification
Key points
  1. On 24 August, Intel presented a 350W data-centre inference card with up to 480GB of memory, alongside client and edge SoCs integrating an NPU rated at up to 17 TOPS.
  2. Qualcomm's Dragonwing AI On-Prem Appliance, described on 20 August, uses a compact 150W form factor. The company says it supports models up to 120B, multi-user inference and RAG, with systems available through several hardware partners.
  3. Xiaomi showed its AI Cube Prototype on 24 August. Additional disclosure on 25 August described 80GB of unified memory, a stated 150W sustained-performance envelope and local deployment of 120B and 3B models, but provided neither pricing nor an availability date.
Signal 01

Data-centre and edge tier: one architecture family spans a 350W inference card and an edge SoC

On 24 August at Hot Chips 2026, Intel introduced three architectures positioned for agents and enterprise workloads. Crescent Island was described as a data-centre GPU for inference in a 350W air-cooled PCIe card, supporting up to 480GB of LPDDR5X memory. Wildcat Lake targets mainstream clients and intelligent edge platforms, integrates an NPU rated at up to 17 TOPS, and combines 2 performance cores with 4 efficiency cores. Diamond Rapids was positioned as a scalable compute foundation and high-performance orchestration processor for enterprise agentic AI.

The announcement spans a data-centre accelerator card, client hardware and an edge SoC; it does not claim that one device completes every workload. Intel's material is primarily an architecture preview and a statement of next-generation product positioning. It does not provide reproducible agent benchmarks, pricing or a complete supply timetable, so claimed gains in throughput, concurrency and efficiency remain vendor targets awaiting independent validation.

What this may mean for enterprise adoption

As model deployment moves towards smaller devices, the likely change is in workload placement rather than the disappearance of the data centre. Sensitive data, low-latency interactions and disconnected operation may be assigned to a local node, while long-context, high-concurrency or complex reasoning may still use a larger compute layer. For an enterprise agent, the meaningful differences may therefore lie in routing across tiers, permissions and auditability, and whether knowledge-graph or ontology context remains consistent at each execution location—not simply in the peak specification of an individual device.

Signal 02

Appliance tier: the same 150W figure does not make two product states equivalent

In an edge-AI briefing published on 20 August, Qualcomm said the Dragonwing AI On-Prem Appliance is based on its Cloud AI 100 Ultra accelerator. The company describes up to 870 TOPS in a compact 150W system, with support for large vision and language models up to 120B, multi-user inference and RAG workloads. Its material says systems can be supplied by hardware partners including Aetina, Advantech and Lanner.

These are vendor specifications and statements of supported workloads, not an independent comparison under common conditions. Qualcomm does not disclose a model, numerical precision, context length, batch size or concurrency configuration that can be aligned with AI Cube. Although both descriptions contain the figures 150W and 120B, those shared figures do not establish equivalent performance, efficiency or software maturity.

What this may mean for enterprise adoption

The enterprise value of a compact local appliance depends on more than its physical size. Compared with a demonstration-only prototype, a partner supply channel and an integration path may shorten the distance from evaluation to deployment. Operating-system and driver maintenance, model catalogues, multi-tenant isolation, observability and the supported service life will determine whether a device can function as shared inference infrastructure. Similar power and parameter figures are only signals about form factor; they cannot replace reliability and cost testing against a real workload.

Signal 03

Desktop prototype tier: AI Cube demonstrates a local combination of 120B and 3B models

At its Xring chip technology event on 24 August, Xiaomi demonstrated an engineering version of the AI Cube Prototype mini PC. ITHome's on-site report says the system combines the Xring O3, O100 and D100 chips, claims a 150W sustained-performance envelope, and supports local deployment of 120B and 3B models with switching between faster and slower systems. Additional information published on 25 August says the prototype has 80GB of unified memory.

The public information remains at prototype and demonstration stage. Xiaomi has not announced a launch date or price. Nor do the available materials identify the 120B model, numerical precision, context length, concurrency setting or measured throughput, so parameter count alone cannot establish response speed, output quality or performance under a sustained workload.

What this may mean for enterprise adoption

The prototype indicates that local execution of large-parameter models is entering the design space of desktop hardware. Unified memory, near-memory bandwidth and thermal limits may become constraints alongside chip compute. Completing a demonstration is not the same as providing an enterprise with a purchasable and maintainable private-deployment node. Driver and model compatibility, remote management, permission isolation, long-term supply and failure recovery will all affect whether this type of device can enter a production environment.

Verification

Sources and verification

  1. Intel Outlines Architectures for Agentic AI at Hot Chips 2026Intel · 2026-08-24 · Official announcement
  2. Edge AI ignites the next industrial revolutionQualcomm · 2026-08-20 · Official announcement

Golden Data has edited this briefing from the public materials listed above. The original sources govern facts and figures. The enterprise relevance sections are Golden Data editorial analysis and do not constitute an endorsement of any third-party product.

← Back to AI Daily Briefing