Edge AI without Compromise: Efficient, Versatile and Accurate Neurocomputing in Resistive Random-Access Memory
AI-generated Key Points
- Edge hardware development is crucial for cloud-level AI functionalities at the edge of the internet
- NeuRRAM is the first multimodal edge AI chip using RRAM CIM
- NeuRRAM delivers high versatility and energy-efficiency 5-8 times better than prior art across various computational bit-precisions
- Inference accuracy is comparable to software models with 4-bit weights on all measured standard AI benchmarks, including impressive results on MNIST image classification, CIFAR-10 image classification, and Google speech command recognition
- NeuRRAM achieved a 70% reduction in image reconstruction error on a Bayesian image recovery task
- These results pave the way towards building highly efficient and reconfigurable edge AI hardware platforms for more demanding and heterogeneous AI applications in the future.
Authors: Weier Wan (Stanford University), Rajkumar Kubendran (University of California San Diego), Clemens Schaefer (University of Notre Dame), S. Burc Eryilmaz (Stanford University), Wenqiang Zhang (Tsinghua University), Dabin Wu (Tsinghua University), Stephen Deiss (University of California San Diego), Priyanka Raina (Stanford University), He Qian (Tsinghua University), Bin Gao (Tsinghua University), Siddharth Joshi (University of Notre Dame), Huaqiang Wu (Tsinghua University), H. -S. Philip Wong (Stanford University), Gert Cauwenberghs (University of California San Diego)
Abstract: Realizing today's cloud-level artificial intelligence functionalities directly on devices distributed at the edge of the internet calls for edge hardware capable of processing multiple modalities of sensory data (e.g. video, audio) at unprecedented energy-efficiency. AI hardware architectures today cannot meet the demand due to a fundamental "memory wall": data movement between separate compute and memory units consumes large energy and incurs long latency. Resistive random-access memory (RRAM) based compute-in-memory (CIM) architectures promise to bring orders of magnitude energy-efficiency improvement by performing computation directly within memory. However, conventional approaches to CIM hardware design limit its functional flexibility necessary for processing diverse AI workloads, and must overcome hardware imperfections that degrade inference accuracy. Such trade-offs between efficiency, versatility and accuracy cannot be addressed by isolated improvements on any single level of the design. By co-optimizing across all hierarchies of the design from algorithms and architecture to circuits and devices, we present NeuRRAM - the first multimodal edge AI chip using RRAM CIM to simultaneously deliver a high degree of versatility for diverse model architectures, record energy-efficiency $5\times$ - $8\times$ better than prior art across various computational bit-precisions, and inference accuracy comparable to software models with 4-bit weights on all measured standard AI benchmarks including accuracy of 99.0% on MNIST and 85.7% on CIFAR-10 image classification, 84.7% accuracy on Google speech command recognition, and a 70% reduction in image reconstruction error on a Bayesian image recovery task. This work paves a way towards building highly efficient and reconfigurable edge AI hardware platforms for the more demanding and heterogeneous AI applications of the future.
Ask questions about this paper to our AI assistant
You can also chat with multiple papers at once here.
Assess the quality of the AI-generated content by voting
Score: 0
Why do we need votes?
Votes are used to determine whether we need to re-run our summarizing tools. If the count reaches -10, our tools can be restarted.
The previous summary was created more than a year ago and can be re-run (if necessary) by clicking on the Run button below.
Similar papers summarized with our AI tools
Navigate through even more similar papers through a
tree representationLook for similar papers (in beta version)
By clicking on the button above, our algorithm will scan all papers in our database to find the closest based on the contents of the full papers and not just on metadata. Please note that it only works for papers that we have generated summaries for and you can rerun it from time to time to get a more accurate result while our database grows.
Disclaimer: The AI-based summarization tool and virtual assistant provided on this website may not always provide accurate and complete summaries or responses. We encourage you to carefully review and evaluate the generated content to ensure its quality and relevance to your needs.