You've seen the economics. This is the architecture behind them — and why every choice here exists to protect that return.你已经看过经济性测算。这里是支撑它的架构——以及每一个设计取舍为何都在守护那份回报。
Compute, power, cooling — each delivered as a factory-integrated module. On site you only connect power, cooling, and network. No on-site system commissioning.算力、电力、冷却——每一项都以工厂集成模块交付。现场只需接入电、冷、网,无需现场系统调试。

The compute engine — pick a liquid-cooled path by workload, or run both on one shared 28–35 °C loop.算力引擎——按负载选择液冷路线,或让两者共用一条 28–35 °C 回路。

A grid-tied power asset — not just backup. 120 min ride-through standard, plus peak-shaving and demand response.并网电力资产——不只是备电。标准 120 分钟穿越能力,兼顾削峰填谷与需求响应。

Free cooling whenever the climate allows, mechanical trim only when it doesn't — across all climates.气候允许时全程自然冷却,不允许时才启动机械补冷——适配各类气候。
AI racks draw 80–150 kW and climb; GPU generations turn over every 1–2 years. The 12–24-month build model can't keep up. The real constraint is no longer how fast you can build — it's how fast you can energize, and how well you work with the grid.AI 机柜功率已达 80–150 kW 并持续上升;GPU 每 1–2 年换代。12–24 个月的土建模式跟不上节奏。真正的约束不再是建得多快,而是通电多快、以及与电网协同得多好。
Where you can build is decided by power. MDCX is engineered to be a grid-cooperative load — so it clears interconnection faster, and earns on power others can't use.能在哪里建,由电力决定。MDCX 被设计成与电网协同的负荷——因此并网审批更快,并能靠别人用不上的电赚钱。
Powered by BESS, not diesel —以 BESS 供电,而非柴发 —
Myth: diesel isn't faster — UPS bridges the gap in both cases. So BESS wins on cost, logistics, and approvals.常见误解:柴发并不更快——两种方案都由 UPS 承担切换间隙。因此 BESS 在成本、物流与审批上均占优。
An AI factory runs ~10× the scale of an enterprise data center. Full 2N / Tier III doubles your infrastructure CapEx on a massive base — straight against the IRR you just modeled.AI 工厂的规模约为企业数据中心的 10 倍。在如此大的基数上做全 2N / Tier III,会让基础设施 CapEx 翻倍——直接冲击你刚测算出的 IRR。
And AI training isn't a financial-transaction workload. It checkpoints, restarts, and fails over at the node level — it carries no per-second SLA penalty. Buying enterprise-grade redundancy for it is capital spent to protect against a risk that doesn't cost you.而且 AI 训练不是金融交易类负载。它有检查点、可重启、在节点级做故障切换——不存在按秒计的 SLA 罚则。为它购买企业级冗余,是花钱防一个本来不会让你损失的风险。
So redundancy is matched to each component's replaceability and blast radius:因此,冗余按各部件的可更换性与故障影响范围来匹配:
| System系统 | Strategy策略 | Logic逻辑 |
|---|---|---|
| Busbar / copper rails母线 / 铜排 | Single-path, high-integrity单路径、高完整性 | Not field-replaceable — redundant paths only add connection points and failure modes.无法现场更换——冗余路径只会增加连接点与故障模式。 |
| Cooling pumps / CDU冷却泵 / CDU | 2N | Supports online maintenance — swap without stopping compute.支持在线维护——更换无需停止算力。 |
| Fans / compressors风机 / 压缩机 | N+1 | Degrade-and-continue on failure, not an immediate stop.故障时降级运行而非立即停机。 |
| Compute nodes计算节点 | Node-level failover节点级故障切换 | Local fault containment — failures don't propagate across the cluster.故障就地隔离——不会在集群内扩散。 |
| Control plane控制平面 | Redundant冗余配置 | No single point of control can take down the whole system.不存在能拖垮整个系统的单点控制。 |
MDCX is a manufactured product, not a construction project: fully integrated and tested at the factory, online on arrival. The first cluster ships in ~4–6 months — and because the design is proven and the supply chain is warm, every batch after lands in 75–90 days.MDCX 是制造出来的产品,不是工程项目:在工厂完成全部集成与测试,到场即可上线。首个集群约 4–6 个月交付——由于设计已验证、供应链处于热状态,后续每批次 75–90 天到位。
CIOS ships with every block: see every sensor, turn every alarm into a ticket, meter every kilowatt-hour and GPU-hour.CIOS 随每套模块交付:看见每一个传感器,把每一条告警变成工单,计量每一度电与每一个 GPU 小时。

Path-addressed telemetry:路径寻址遥测: sgp01.pod002.cdu000.fws.supply.flow

Alarms → tickets → SLA, policy-gated setpoints.告警 → 工单 → SLA,设定点受策略约束。

Usage metering, capacity headroom, ops reports.用量计量、容量余量、运维报表。
Fast deployment and lean redundancy don't mean cutting corners. Every block is built from tier-one industrial hardware and certified to recognized standards.快速部署与精简冗余不等于偷工减料。每套模块都采用一线工业级硬件,并按公认标准取得认证。

No compromise components. Every block is assembled from proven, tier-one industrial hardware.部件不做妥协。每套模块均由经过验证的一线工业级硬件装配而成。

All components meet UL / CE / CSA and the applicable regional standards.所有部件符合 UL / CE / CSA 及适用的地区标准。

End-to-end system-level certification targeted for completion in 2027.端到端系统级认证目标于 2027 年完成。
Two container profiles — chosen by the GPU you want to run and the business you're building. Frontier training on Blackwell, or a lean inference business at scale.两种集装箱配置——取决于你要跑的 GPU 与你要做的生意。基于 Blackwell 的前沿训练,或规模化的精简推理业务。

Maximum-density training & high-performance inference. Premium, frontier-grade compute — high CapEx, highest return per container.极限密度训练与高性能推理。前沿级高端算力——CapEx 高,单箱回报最高。

Pure inference at scale. Lower entry cost, simple operations, durable continuous revenue.规模化纯推理。入门成本更低、运维简单、收入持续稳定。
Running any of the above →运行上述任一方案 → CIOS · included →CIOS · 已包含 →