Publications

Our teams aspire to make discoveries that impact everyone, and core to our approach is sharing our research and tools to fuel progress in the field.

people standing in front of a screen with images and a chipboard

Our teams aspire to make discoveries that impact everyone, and core to our approach is sharing our research and tools to fuel progress in the field.

Sort By
  • Title
  • Title, descending
  • Year
  • Year, descending
1 - 15 of 172 publications
Preview abstract Every abstraction layer in the modern software stack exists to solve legitimate problems—coordinating independent developers and enforcing trust across boundaries. However, each layer exacts a Cognitive Tax: overhead paid not for correctness, but for human coordination. Published measurements bound this non-computational overhead at ~59% across the ISA frontend, ABI, IEEE 754 logic, and per-die guard-band slack. We propose a shift from shipping static artifacts to distributing formal intent. Developer intent is translated into Z3 invariant bundles, which a local Neural-Symbolic Oracle then synthesizes into `Asemantic Code'—an artifact governed by load-time proof certificates and mathematically mutated to exploit the specific manufacturing physics of its execution die. View details
Preview abstract While MVP and CLEAN Architecture are popular Android patterns, they often introduce boilerplate or lack reactivity. This article introduces the Reactive Data Layer Architecture (RDLA), an offline-first, push-based data layer pattern designed for modern Android apps using Jetpack Compose and Room. Using a heart rate tracking example, we demonstrate how RDLA provides robust local-remote synchronization, clean separation of concerns without Use Case overhead, and simplified unit testing via the TestExtensions pattern. View details
Physical Design Aware Verification Methodology for Closing Coverage Gaps in SharedBus MBIST
Shivam Tulsyan
Vasudevan Pillai A
Maheedhar Jalasutram
Prachi Sinha
Mayank Parasrampuria
2026
Preview abstract The industry shift toward SharedBus MBIST architectures has successfully mitigated the Power, Performance, and Area (PPA) bottlenecks associated with traditional embedded memory testing. However, reusing functional paths for testing introduces severe verification challenges, as conventional MBIST algorithms often fail to detect intricate mapping errors like data-bus scrambling, tiedoff data bits, and irregular address bits decoding. If left undetected, these discrepancies in implementation result in silent coverage gaps and ineffective memory repair mechanisms. This paper proposes a robust assertion-based RTL verification methodology specifically designed to close these gaps in SharedBus MBIST implementations. By deploying a Walking-0 pattern and continuous monitors across SharedBus and physical memory interfaces, the methodology enforces a strict set of verification rules. Experimental results validate this approach, demonstrating the successful identification of critical implementation bugs across multiple vendor cores that escaped conventional verification. The paper concludes by proving that the overhead of this methodology is minimal and highly justified by the resulting improvements in silicon quality. View details
Preview abstract Lightweight execution environments like the Little Kernel (LK) are commonly deployed in post-silicon validation to assess software-hardware interactions. These bare-metal kernels, however, lack the sophisticated power management features present in full operating systems, such as the Generic Power Domain (GenPD) framework. Instead of building complex software abstractions that simulate production-grade power management drivers, this paper applies a Design Verification (DV) approach to post silicon. By discarding standard software paradigms, the introduced framework leverages fundamental bare-metal kernel primitives to intentionally engineer synthetic, non-deterministic stressors. Eliminating intermediate software layers allows us to utilize the bare-metal setup to drive chaotic, aggressive stimuli straight to the hardware execution layer. View details
DDRop: Generic Memory Interposer Attacks on Confidential VMs by Dropping DDR5 Writes
Jesse Demeulemeester
Stefan Gloor
Patrick Jattke
David Oswald
Martin Thompson
Kaveh Razavi
Ingrid Verbauwhede
Jo Van Bulck
ACM Conference on Computer and Communications Security (CCS) (2026)
Preview abstract Trusted Execution Environments (TEEs) are increasingly deployed in the cloud to protect sensitive workloads through hardwareenforced isolation, remote attestation, and transparent memory encryption. However, to meet memory performance and size demands, modern TEEs omit cryptographic freshness guarantees, leaving them vulnerable to replay attacks by adversaries with physical memory access. Prior work demonstrated low-cost active interposition attacks on DDR4 without requiring expensive specialized equipment, but these techniques do not extend to DDR5, where existing approaches are limited to passive ciphertext side-channel analysis and rely on bus downclocking to accommodate legacy memory bus analyzers. We present the first low-cost (<200$) DDR5 RDIMM interposer capable of active fault injection at native speeds. By injecting targeted parity errors to silently discard cache line writebacks, we introduce DDRop, a new primitive that exploits the absence of cryptographic freshness to break the integrity of Intel TDX, Scalable SGX, and AMD SEV-SNP. Building on this primitive and targeting the APIs exposed by the TDX module and AMD Secure Processor, we show that adversaries can gain ciphertext access, copy arbitrary victim pages, and inject malicious secure page-table entries. We demonstrate end-to-end attacks on an up-to-date TDX platform, including forcing any TD into debug mode and forging attestation reports. While software-level mitigations, including timing-based interposer detection and API hardening, may reduce the attack surface, our results demonstrate that active DDR5 bus interposition is practical at low cost, highlighting the need for robust cryptographic memory integrity protections against physical adversaries. View details
Modeling multi-chiplet architectures for TPU co-design
Pritha Doddahosahally Narayanappa
Hung-Ming Hsu
Khai Tran
Hardie Cate
Narges Shahidi
Zhijie Deng
Avinash Lingamneni
Thejasvi Vijayaraj
Aditya Yanamandra
Sameer Kumar
Ming Liu
Lluis-Miquel Munguia
2026
Preview abstract We present an experience report on modeling multi-chiplet hardware accelerators for TPU co-design, and argue that detailed chiplet modeling is necessary to identify bottlenecks and optimizations for future hardware design. Our chiplet modeling tool (CMT) shows that at lower target serving latencies, serving queries/second/chip (QPS/chip) can be up to $2.2\times$ different with chiplet modeling compared to modeling with aggregated compute and memory resources. With this tool, we identify interdependencies between chiplet modeling and the rest of the system: we highlight an example where cost savings from prefetching weights from HBM to local SRAM need to be incorporated in the selection heuristic for chiplet sharding strategies. Using CMT, we explore performance sensitivity to chiplet properties in a design space exploration that highlights considerations for high-performance serving. We use an open-source MoE model (OSS-MoE) served on Ironwood-like TPU architectures as a case study. We conclude by discussing the need to establish best sharding practices to narrow this large design search space. View details
Preview abstract To meet aggressive time-to-market goals, modern mobile SoC architectures require the concurrent development of custom Compute-IPs and surrounding subsystem integration logic (1PIPs and 3PIPs). However, this parallel execution creates a critical verification deadlock: the subsystem cannot be validated until both the Compute-IP and volatile, in-flight 1PIPs reach physical RTL maturity. Consequently, "integration-killer" bugs—such as protocol handshaking deadlocks and clock/reset sequencing mismatches—remain hidden until late in the design cycle when RTL rework costs are prohibitive. To break this bottleneck, we present a verification-driven methodology utilizing a silicon-proven Golden Proxy, Direct-Execution Traffic Profiles, Programmable Sequencers and Automated Protocol Converters. This framework completely decouples parallel hardware dependencies, pre-pulling critical inter-IP mismatch discoveries months ahead of traditional integration milestones. View details
Preview abstract Modern compute subsystems have evolved into complex architectures where custom Compute-IPs interface directly or downstream with a mix of in-flight 1st-party IPs (1PIP), stable legacy components, and pre-verified 3rd-party IPs (3PIP). These connections—whether coherent, non-coherent / configuration, for custom use-cases utilizing in-house protocols, etc—often require concurrent development of both the compute tiles and their integration logic to meet aggressive time-to-market goals. However, this parallel approach introduces a critical bottleneck: the subsystem cannot be verified until both the Compute-IP and the 1PIPs reach maturity. Consequently, fundamental functional misalignments—such as protocol handshaking deadlocks, clock and reset sequencing issues, and architectural assumption mismatches—often remain hidden during IP development phases, resulting in a high-risk discovery tail where "integration-killer" bugs are uncovered only when RTL rework costs and schedule impacts are prohibitive. To break this deadlock, we present a verification-driven methodology that utilizes a silicon-proven Golden Proxy, Direct-Execution Traffic Profiles, and Programmable Sequencers to provide a functional proof of concept and pre-pull critical inter-IP interaction mismatch discoveries months prior to traditional integration milestones. View details
Preview abstract SoC Flat IR/EM signoff is generally done for multiple cycles thus generally mandating 2+ days to cover a single scenario- not only is the coverage limited but also expensive since even small ECO fixes trigger full analysis repeat (of same resources). Additionally, designs which have multiple hierarchical block level instantiations - this is massively computationally redundant. Reduced Order Model (ROM) Flow: Hierarchical Abstraction for IR/EM Signoff: ROM eliminates computational redundancy by using abstract representations of pre-verified blocks. It leverages tweaked SoC flat analysis to have appropriate block level details to enable 10-20× faster SoC turnaround and broader scenario coverage. Below is its mechanism: The Common Connection Layer (CCL) acts as the electrical boundary between block & SoC top. ROM preserves full detail only at the CCL and CCL-1, while the lower metal layers (M0 to CCL-2) are rolled up into equivalent impedance model to maintain signoff accuracy. Designers use a mix of detailed instances for same critical block with reduced instances to optimize resource usage as shown in Fig 1. The Validation Problem with ROM- Trust Gap: Context Mismatch: ROMs are generated in standalone conditions, failing to account for top-level grid impedance and adjacent block coupling. Fidelity & Coverage Loss: Abstracting 12-14 layers can mask local voltage violations; current manual spot-checks are insufficient since these fail to quantify if CCL node voltages in all ROM instances match their power-domain & scenario specific simulation values Objective of this work: Systematic validation across all ROM instances & all power domains in a quick (wall time ~mins for SoC) else it would offset ROM runtime benefits. Quantitative fidelity metrics with low violation thresholds & spatial coverage for debug to understand root cause of localised errors. View details
Preview abstract Limitations in Sign-off Methodology: Traditional STA corner selection 10% lower STA corner from PMIC voltage is selected- design is constantly optimised for 10% lower voltage, thereby failing to build margin against differential drop. IR aware STA does not account timing path’s geometric imbalances (logic depths), net dominated interconnect skews (net delays & metal layer variation) - all dominant in advanced process nodes. Furthermore, this is workload dependent: fixing IR STA violations does not build margins on unseen vectors. Frequent Silicon issues due to these gaps: Low voltage mode scan shift Vmin jumps need to meet slack on paths which become exponentially sensitive to voltage gradients. Even small differential IR drops (capture & launch traversing through contrasting IR hotspot & cool regions) cause catastrophic slack loss High divergence paths with structural imbalances-where clock paths are net-dominated & are operated at high speeds often fail to meet hold timing, despite good pre-silicon margins due to high interlayer metal-sheet & via resistances in lower process nodes. Additionally, there is considerable PPA impact -higher dynamic & leakage power in clock & data paths respectively in divergent paths. Our proposed solution aims to address above gaps. View details
TCO-driven Storage Provisioning for Exascale Data Centers
Timothy Kim
Prashant Nema
Jai Menon
Rashmi Vinayak
Gregory R. Ganger
2026
Preview abstract Recent changes in data temperatures and storage device characteristics, both mechanical disk drives (HDDs) and solidstate drives (SSDs), expand the set of deployment options for exascale storage. Until recently, exascale storage systems followed a pattern of placing most data on HDDs with smaller amounts of SSD storage used for caching and performance-critical workloads. Exascale storage provisioning and dataset placement trade-offs have now changed. This paper describes a total cost of ownership (TCO) model that captures primary aspects of modern deployments and uses it to explore the new trade-off space. Using capacity and performance telemetry information for 43 production datasets+workloads at two large hyperscalers, we show significant changes from prior analyses of workloads and storage placement decisions across a multitude of storage device types. We also introduce a storage cluster TCO optimizer that identifies the lowest-TCO grouping and assignment of datasets to device types, exposing a number of insights that can help guide future deployments. For example, our analysis shows that the highest-density SSDs are particularly favorable for clusters with heavy AI/ML workloads but are only cost-effective at exascale when combined with high-density HDDs. Finally, we use our framework to evaluate how storage provisioning and overall TCO change as a function of key parameters like device write amplification, cluster power bounds, and the maximum number of device types allowed. View details
Preview abstract This paper provides an overview of Google's TPUs across five generations, from TPU v2 to Ironwood, highlighting their evolution as scalable, resilient, and sustainable supercomputers for AI training. It details the TPU’s stable architecture and microarchitecture, which has surprisingly easily accommodated the rapidly changing deep neural network workloads, such as the rise of Transformers. Key advancements over eight years include 10x increase in HBM capacity and bandwidth per node, a 100x increase in peak node performance, and a 3600x increase in supercomputer performance. The paper also discusses the role of optical circuit switches and built-in self test in enhancing resilience, how TPU’s carbon footprint was reduced by improving embodied carbon emissions per floating point operation and a 30x gain in performance per Watt. It concludes by identifying six features that may well characterize the successful AI accelerators of this decade. View details
Enhancements in Memory Test for Improved Diagnosability and Comprehensive Multi-Bank Test
Shivam Tulsyan
Veerabhadra Rao Vasa
Prachi Sinha
Mayank Parasrampuria
2026
Preview abstract The relentless scaling of System-on-Chip (SoC) architectures toward sub-5nm nodes has driven a shift from standard compiled memories to highly optimized custom memory designs to meet aggressive Power, Performance, and Area (PPA) targets. However, these custom layouts introduce unique physical defect mechanisms and layout-dependent coupling faults that are often invisible to industry-standard March tests. This paper proposes a novel custom testing methodology specifically engineered for multi-bank custom memory architectures in modern SoCs. By implementing a programmable, bank-aware Built-In Self-Test (BIST) architecture, the proposed solution enables concurrent multi-bank stressing to detect inter-bank crosstalk while maintaining strict power-density limits through a custom scheduling algorithm. This methodology provides a scalable framework for ensuring high reliability in performance-critical silicon environments. Industry standard memory built in self test does not cover this. This paper presents a novel custom memory testing methodology utilizing Soft Programmable Opset Logic Addition to address potential coupling-related faults during simultaneous multi-bank access in multi-port memories. While the proposed multi-bank BIST identifies the presence of inter-bank coupling, physical localization of these defects in sub-5nm nodes requires non-invasive back-side analysis. Our methodology includes a 'Diagnostic Mode' that allows the BIST to loop specific stress patterns, enabling Laser Voltage Probing (LVP) to capture high-resolution internal waveforms. This synergy between custom BIST and optical probing significantly reduces the Time-to-Yield (TTY) by pinpointing the exact physical origin of marginal delay faults. View details
Preview abstract System coherency verification is vital for ensuring data consistency in complex memory hierarchies, but late integration often delays bug discovery. This paper presents a "left-shift" approach to accelerate coherency verification. We detail three key aspects: early verification via a stitched DUT for initial testing; automated stimulus generation using third-party tools like Cadence Perspec to cover complex system-level scenarios; and automated checker generation, complemented by an in-house tool (DICE). This methodology significantly reduces test/checker development time, enables faster test creation for corner cases, and results in better system-level coherency coverage, finding critical bugs earlier in the design cycle. View details
×