A startup most readers have never heard of published a whitepaper this month that, if any of it holds up, points at how the post-Nvidia AI chip world could actually be built. PhantaField, a small semiconductor company with no shipped silicon on record, claims its Sophon PFG-1 design hits 2.10 PB/s of in-tile weight bandwidth - roughly 95x what an HBM4-class Nvidia package delivers, per the same paper - by stacking 330 GB of DRAM directly on top of the logic. No HBM. No off-package memory. At a claimed bill-of-materials of $8,358 per die.
The honest framing: this is a paper architecture, not a product. There is no evidence the company has taped out, fabricated, or shipped silicon, and the load-bearing physics claims sit on academic references rather than measured hardware. But the design choices are concrete, the comparison table is detailed, and the directions line up with where the wider chip industry has been pointing. Local-AI readers should know about it.
What PhantaField Actually Claims
The whitepaper (Revision 4.1, dated June 2026) describes a 750 mm² die built on a 28 nm bulk-Si base topped by 64 alternating tiers: 32 tiers of 2D transition-metal-dichalcogenide (MoS₂ and WSe₂) logic and 32 tiers of 2T0C gain-cell DRAM. That gives roughly 110 Mb/mm² of memory density, 1.8-second data retention at room temperature, and a refresh power of about 0.08 W across the full die - numbers that, if reproducible, would eliminate the dominant cost and power draw on modern AI accelerators: stacked high-bandwidth memory.
Compute sits in 131,072 compute-in-memory (CIM) tiles, each a 256×256 weight subarray with a binary sense amp and an 8-level adder tree. Throughput clocks in at 2,100 TFLOPS BF16, 4,200 TFLOPS FP8, and 8,400 TOPS INT8. The numbers most self-hosters will care about: 14,438 tokens per second on 80B FP8 decode at batch size one, jumping to 72,188 tok/s with INT4 plus speculative decoding, at 25.8 mJ per token. For an 80B BF16 train at batch one, the design claims 2,406 tok/s at 564 W average draw; the whitepaper’s own comparison table puts the same workload on a Rubin R200 die at roughly 1,800 W TDP, without quoting a Nvidia-side tokens-per-second figure.
The BOM comparison is where the proposal stops looking like an incremental improvement. PhantaField puts the Sophon PFG-1 die at roughly $8,358 in materials, versus an estimated $82,800 for Rubin R200 and $96,700 for AMD’s MI455X. The three-year TCO for 80B inference lands at $8,807. These figures are paper-stage and depend on yield assumptions the company does not yet defend with measured data.
Why The Memory Hierarchy Is Where The Real Fight Is
This story is not about a single upstart. It is about which direction the industry runs as Nvidia’s HBM-class memory hierarchy becomes the binding constraint on inference economics. HBM4 - the standard memory riding shotgun with Rubin and MI455X in 2026 - is in a multi-year supply crunch that has lifted Micron’s market cap past Meta and Tesla on bets that demand outruns supply through 2027. (The Nvidia page itself is the corporate Rubin platform announcement, which puts Rubin in full production for H2 2026 but does not publish a per-die 80B token-per-second number.) Memory, not raw FLOPs, sets the ceiling on how cheaply a self-hoster can serve a frontier-scale model.
That pressure has driven an entire category of bets on inference-focused accelerators with non-HBM memory: wafer-scale designs at Cerebras, in-memory compute at Mythic and Syntiant, SRAM-heavy inference dies from Groq, and a wave of custom-silicon partnerships with Broadcom from Anthropic, OpenAI, and Meta. The PhantaField paper sits in that same lane, but takes a more aggressive swing at the memory wall by replacing HBM with monolithic-3D gain-cell DRAM. If the production physics hold - 1 fA/µm off-current at the device level, ALD growth temperatures under 450 °C, and 90 nm-pitch monolithic inter-tier vias at yield - then the bandwidth ceiling moves by two orders of magnitude. If they do not, the paper joins a long line of architectural designs that were great to think with and never made it to silicon.
What To Make Of The Claims
The honest read is that PhantaField is publishing in a venue that other researchers can scrutinize, with a comparison table that lets reviewers check the math against published Nvidia and AMD specs. The energy-per-token and bandwidth-per-dollar numbers are internally consistent. The physics, however, are the load-bearing wall. Multi-second retention in a 2T0C gain cell built on a 28 nm Si base with sub-100 nm monolithic 3D vias is not yet demonstrated at production scale anywhere in the public literature the paper cites. The 64-tier BEOL 2D-TMD stack and the wafer-scale ALD growth it implies have no foundry precedent in the references.
PhantaField is also a small company with no public foundry agreement and no disclosed customer. The whitepaper cites only academic references and a Morgan Stanley teardown estimate for the HBM comparisons - no customer letters, no taped-out parts, no conference tape-out announcements. Readers should treat the document as a serious architectural proposal worth understanding, not as a roadmap for hardware arriving in 2027.
What This Means
For self-hosters, the practical takeaway is not “buy a Sophon.” The takeaway is that the next round of inference economics will be settled by memory hierarchies, not by raw FLOP counts, and that the HBM oligopoly is no longer a stable target. A handful of credible teams are now publicly attacking it from different angles. PhantaField is one of the bolder bets on a publishable whitepaper.
For readers who track the local-AI beat, the next twelve months will tell which of these bets actually tape out. Watch for any PhantaField foundry announcement, any production yield update on monolithic 3D stacks from TSMC or imec, and any second source of memory-direct-on-die designs from the hyperscalers’ custom-silicon partners.
The Bottom Line
PhantaField’s PFG-1 whitepaper is a serious architectural proposal from a small company with no shipped silicon. If 5% of the bandwidth and BOM claims survive contact with a foundry, the Nvidia-HBM duopoly has a real challenger on the horizon.