AI Infrastructure
Photoreal flythrough of a circuit board, its components lit like a city at night

Aditya Morey · NVIDIA · AI Hardware DFX

Silicon observabilityfor the AI compute stack.

Debug and observability systems that turn hidden AI/HPC silicon state into repeatable root-cause workflows.

Scan paths reach what the package hides.

JTAG/IJTAG, MBIST, IO BIST: engineered access into otherwise invisible silicon state.

Capture, replay, root cause.

Deterministic debug platforms shared across teams, not one-off scripts.

Observability designed in at RTL.

Built into the silicon up front, not bolted on after the bug.

In one line

I build and study debug, verification, and observability systems for complex compute, turning low-level silicon behavior into repeatable, inspectable workflows engineers can actually use.

Portrait of Aditya Morey
Aditya MoreyNVIDIA · LCI paper (SSRN) · debug-platform work

Focus areas

Cross-team debug toolingAI/HPC silicon systemsScan/JTAG state capturePre-silicon verificationPlatform observability
Affiliations & proof
NVIDIAPreviously · IntelSSRNAdditive-manufacturing patent
Multi-$Msilicon test-cost savings~20%yield increase~15%spec-validation efficiency

Latest writing

Manifesto

Make every layer of the system inspectable.

Or call it luck.

Frontier AI silicon and frontier AI agents fail the same way: in the dark, between layers, under pressure. Observability is not a dashboard. It’s the thing that turns those failures into root-cause workflows.

§ 01

Silicon state capture.

Scan paths, JTAG/IJTAG, MBIST, IO BIST: engineered access into otherwise invisible behaviour.

§ 02

Repeatable debug platforms.

Tooling shared across teams: deterministic capture, replay, and root-cause workflows.

§ 03

Agent observability.

Tracing, escalation checkpoints, replay: the same primitives, applied to autonomous systems.

§ 01

Start here

A focused hub for AI hardware infrastructure.

Cover of Designed to Be Debugged

New · book · first research edition

Designed to Be Debugged

Capability without engineered evidence is not deployment readiness. Why autonomous AI needs state capture, replay, and escalation before it can be trusted: silicon debug discipline, translated for agentic systems.

Publication hub

The Observability Stack

A curated technical publication surface for essays, frameworks, and labs on observability for complex compute systems.

Open Stack

Who I talk to

Three ways in: start where you are.

§ 02

Public artifacts

Working demos, papers, and essays that make the technical thesis inspectable.

demos / papers / essays / inspectable

Lab

Scan Chain / TAP Visualizer

Interactive visualizer for scan-chain state movement, TAP-style control flow, and state-capture intuition.

Open
Lab

Fault Injection / Observability Demo

Public demo showing how observability changes fault isolation and root-cause workflow.

Open
Paper / framework

The Cost of Usable Intelligence

Published framework for reasoning about useful AI output under latency, reliability, safety, and quality constraints.

Open
Technical essay

Observability is an underbuilt safety primitive

Flagship essay on tracing, replaying, and debugging agent behavior before failures scale.

Open
§ 03

Featured work

Making complex hardware observable.

Selected professional themes, described without employer-confidential details. The strongest thread is platform leverage: turning low-level silicon behavior into repeatable workflows engineers can use.

Complex AI/HPC
silicon system
Freeze
Capture
Decode
Analyze
Root
cause

repeatable debug workflow / platform observability / silicon readiness

Debug platforms for deterministic state capture

Built and maintain internal scan-debug tooling used by multiple engineering teams to support deterministic state capture and root-cause workflows.

Cross-functional silicon readiness

Work across DFX, design, verification, bring-up, and debug stakeholders to improve observability, readiness, and platform reliability for next-generation AI/HPC silicon programs.

Hardware/software debug boundary

Focus on making complex hardware systems observable, debuggable, and reliable through state capture, test access, and engineering productivity tooling.

§ 04

Current focus

Core practice and emerging research arcs.

Core practice means direct professional or published work. Emerging research arcs means active study, writing, and future work.

Core practice

Emerging research arcs

Silicon observability and debug platforms

State capture, scan access, JTAG/TDI-based workflows, MBIST context, and platform readiness.

Autonomous agent observability and red-team evals

Failure modes, telemetry, containment, auditability, and cyberphysical risk evaluation.

AI infrastructure economics

Useful output cost, latency, reliability, safety constraints, and infrastructure productivity.

Quantum-classical infrastructure

Control planes, accelerated classical co-processing, observability, and integration bottlenecks.

Contact

For conversations around AI hardware observability, frontier infrastructure, research evals, or high-leverage technical systems.