The Future Stack: Powering AI Beyond Moore’s Law in 2026

0
4

Key Takeaways

  • The price of completing a qualified AI task at constant performance has dropped roughly tenfold each year from 2021‑2024 – a 1,000‑fold improvement in three years, according to Andreessen Horowitz; MIT researchers place the recent frontier rate at 5‑10× per year.
  • This deflationary pace exceeds historical transistor scaling (Moore’s Law) and stems from multiple, compounding curves – chip architecture, system design, model efficiency, software, security and deployment – rather than a single silicon curve.
  • As the cost per qualified outcome falls, more AI workflows become economically viable, expanding opportunities across infrastructure, software, security, autonomous systems, personal agents and real‑world embodied AI.
  • Index‑based vehicles such as the ROBO Global Artificial Intelligence Index (THNQ) and its associated ETFs capture the full‑stack advances – hardware, system architecture, models, data, security and deployment – that together drive the next phase of the AI market.

Why Moore’s Law No Longer Captures AI Progress
For decades, Moore’s Law served as the “definitive, tried and true method for describing the rate of change in computing, from a technical perspective and on a cost and performance curve.” It shaped the modern world by delivering predictable transistor density gains. AI, however, does not map onto a single curve; its advancement is a “much more complicated and interconnected set of variables.” Leading‑edge chips still improve, but performance now also derives from specialized architectures, memory hierarchy, networking and the software stack above silicon. As the article notes, “Silicon improves in steady generational steps. Model architecture improves in jumps, when someone changes how the software works. A single blended rate averages across both describes neither.” The more appropriate lens is the price of a completed task that meets the required standard, because it folds together hardware efficiency, model design, serving software, data, security and human review into a single outcome‑focused metric.


The Full‑Stack Race: Hardware‑Software Co‑design
Recent announcements illustrate how system‑level innovation adds fresh efficiency curves beyond raw transistor scaling. At AMD’s Advancing AI 2026 event, AMD and Cerebras unveiled a joint inference platform that pairs AMD’s Helios processor with the Cerebras Wafer‑Scale Engine. AMD handles prompt processing and high‑volume throughput while Cerebras generates the answer, yielding “up to five times more tokens per second per watt than a Cerebras‑only configuration.” This demonstrates that system architecture can add another efficiency curve.

On the software side, China‑based Moonshot AI released the open‑weight model Kimi K3, which activates only 16 of its 896 specialist components per token. Moonshot estimates the design improves overall scaling efficiency by about 2.5 × over its predecessor, Kimi K2, showing that model sparsity and dynamic routing can drive performance gains independent of silicon.

Security and trust are also being woven into the stack. Nvidia launched the Open Secure AI Alliance, uniting cloud providers, cybersecurity firms, enterprise software vendors and AI researchers to treat “identity, permissions, guardrails, logs, and evaluation as part of agent security.” The alliance extends Nvidia’s earlier Nemotron Coalition, which pools research, data, evaluations and computation across labs, underscoring that agent safety is now a full‑stack concern.


From Token Prices to Qualified Outcomes
Focusing solely on the cost per token can be misleading. As the article warns, “Cheap tokens does not necessarily mean cheap AI or a good outcome.” A low‑price model may require many retry attempts or extensive human review, whereas a higher‑cost model might finish the same task correctly in a single pass.

Deterministic tasks—such as verifying that code compiles or that a required field is present—lend themselves to straightforward cost‑per‑token analysis. Non‑deterministic work, which involves chains of dependencies, multiple parties and real‑world interaction, may produce variable outputs; “reasonable people may prefer different outcomes.” The modern economy resolves this ambiguity through standards, review and accountability.

Artificial Analysis provides a useful benchmark: its index tracks agentic work, coding, scientific reasoning and general knowledge, reporting average cost per benchmark task. For example, “Claude Opus 5 scores 61 at $2.03 per task. DeepSeek V4 Flash scores 44 at four cents.” Here, a lower‑scoring model can be the economic choice if it satisfies the required quality bar, creating a deflationary impact on the cost of “intelligence.”

Nevertheless, the metric has limits. Benchmarks grade work inside controlled environments, while valuable business processes often span software systems, organizations, the physical world and subjective judgment. As the piece observes, “Cost per task is a useful frontier, but not a universal price tag.”

Industry leaders are converging on a more nuanced view. OpenAI frames cost per successful task as a function of price, compute and the likelihood of reaching the correct result. Nvidia uses “intelligence per dollar” to refine models post‑training, explicitly pairing capability and quality with cost.


How Falling Cost per Qualified Task Expands Opportunity
When the expense of achieving a qualified AI outcome declines, previously prohibitive workflows become economical to automate or augment. Agents may consume more tokens as they plan, invoke tools and recover from errors, extending AI’s reach into the physical realm—summoning services, interacting on behalf of individuals or organizations, or acting as embodied agents in real life. This amplification of scope, with broader access to systems and data, heightens cybersecurity demands; autonomous agents require observability at scales that dwarf prior monitoring needs.

The opportunity stretches from silicon and cloud infrastructure through networking, security, data platforms, business processes and industry‑specific applications. Nebius Group (NBIS), an AI‑cloud provider, exemplifies the enabling layer: its Nvidia partnership spans AI‑factory design, inference software, hardware deployment and fleet management, and it was an early launch partner for Kimi K3, showing how ecosystem benefits accrue when open‑source models succeed.

The ROBO Global Artificial Intelligence Index (THNQ) captures these multilayered advances. Unlike the Moore’s Law era, where investors watched a single transistor‑density curve, AI demands vigilance over how chips, system architecture, models, data, security and deployment improve together. Companies that lower the cost of useful work—or expand what AI can do—are constructing the next phase of the market. THNQ underlies the ROBO Global Artificial Intelligence ETF (THNQ) and the L&G Artificial Intelligence UCITS ETF (AIAI.LN), offering a tradable basket for those seeking exposure to this full‑stack evolution.


Outlook
The trajectory suggests that the “price of intelligence” will continue to fall as multiple innovation curves compound. Investors and technologists alike must shift from tracking a lone silicon metric to evaluating the integrated performance‑cost frontier that defines useful AI work. As the cost per qualified task drops, we can expect a broader diffusion of AI agents across enterprise, consumer and industrial domains, accompanied by rising demands for robust security, observability and responsible deployment—hallmarks of the next wave of AI‑driven economic transformation.

https://www.etftrends.com/artificial-intelligence-content-hub/beyond-moores-law-full-stack-driving-ai-2026/

SignUpSignUp form

LEAVE A REPLY

Please enter your comment!
Please enter your name here