TorchAO: PyTorch-Native Training-to-Serving Model Optimization

AI-generated keywords: TorchAO PyTorch-Native Model Optimization Large Language Models (LLMs) Quantization

AI-generated Key Points

TorchAO is a model optimization framework integrated into the lifecycle of large language models (LLMs).
It offers an end-to-end workflow for AI models by leveraging PyTorch-native techniques.
Supports various model optimization methods such as FP8 training, quantization-aware training (QAT), post-training quantization (PTQ), and 2:4 sparsity across different backends.
Integration with pre-training tools like TorchTitan and fine-tuning platforms like TorchTune and Axolotl ensures a unified workflow from development to deployment in serving environments.
Facilitated recent launches of quantized Llama 3.21B/3B and LlamaGuard3-8B models by bridging gaps in model optimization pipelines.
Offers flexibility in targeting diverse hardware environments while enabling efficient model optimization techniques for enhanced performance.
Open-source project available at https://github.com/pytorch/ao.

Also access our AI generated: Comprehensive summary, Lay summary, Blog-like article; or ask questions about this paper to our AI assistant.

Authors: Andrew Or, Apurva Jain, Daniel Vega-Myhre, Jesse Cai, Charles David Hernandez, Zhenrui Zheng, Driss Guessous, Vasiliy Kuznetsov, Christian Puhrsch, Mark Saroufim, Supriya Rao, Thien Tran, Aleksandar Samardžić

ICML 2025 Workshop on Championing Open-source DEvelopment (CODEML 2025)

arXiv: 2507.16099v1 - DOI (cs.LG)

5 pages, 3 figures, published in CODEML@ICML25

License: CC BY 4.0

Abstract: We present TorchAO, a PyTorch-native model optimization framework leveraging quantization and sparsity to provide an end-to-end, training-to-serving workflow for AI models. TorchAO supports a variety of popular model optimization techniques, including FP8 quantized training, quantization-aware training (QAT), post-training quantization (PTQ), and 2:4 sparsity, and leverages a novel tensor subclass abstraction to represent a variety of widely-used, backend agnostic low precision data types, including INT4, INT8, FP8, MXFP4, MXFP6, and MXFP8. TorchAO integrates closely with the broader ecosystem at each step of the model optimization pipeline, from pre-training (TorchTitan) to fine-tuning (TorchTune, Axolotl) to serving (HuggingFace, vLLM, SGLang, ExecuTorch), connecting an otherwise fragmented space in a single, unified workflow. TorchAO has enabled recent launches of the quantized Llama 3.2 1B/3B and LlamaGuard3-8B models and is open-source at https://github.com/pytorch/ao/.

Submitted to arXiv on 21 Jul. 2025

Ask questions about this paper to our AI assistant

You can also chat with multiple papers at once here.

AI assistant instructions?

Results of the summarizing process for the arXiv paper: 2507.16099v1

Comprehensive Summary
Key points
Layman's Summary
Blog article

In their paper titled "TorchAO: PyTorch-Native Training-to-Serving Model Optimization," authors Andrew Or, Apurva Jain, Daniel Vega-Myhre, Jesse Cai, Charles David Hernandez, Zhenrui Zheng, Driss Guessous, Vasiliy Kuznetsov, Christian Puhrsch, Mark Saroufim, Supriya Rao, Thien Tran, and Aleksandar Samardžić introduce TorchAO as a model optimization framework integrated seamlessly into the lifecycle of large language models (LLMs). The framework leverages and techniques to offer an end-to-end workflow for AI models. It supports various model optimization methods such as FP8 training, quantization-aware training (QAT), post-training quantization (PTQ), and 2:4 sparsity across different backends including server CPU/GPU and mobile/edge devices. The authors emphasize the importance of TorchAO's integration with pre-training tools like TorchTitan and fine-tuning platforms like TorchTune and Axolotl in ensuring a unified workflow from model development to deployment in serving environments supported by frameworks like HuggingFace, vLLM, SGLang,and ExecuTorch. By bridging gaps in the fragmented space of model optimization pipelines,TorchAO has facilitated recent launches of quantized Llama 3.2 1B/3B and LlamaGuard3-8B models. The paper concludes by inviting contributions to the open-source project at https://github.com/pytorch/ao. The authors highlight the flexibility of in targeting diverse hardware environments while enabling efficient model optimization techniques for enhanced performance across various applications.

- TorchAO is a model optimization framework integrated into the lifecycle of large language models (LLMs).
- It offers an end-to-end workflow for AI models by leveraging PyTorch-native techniques.
- Supports various model optimization methods such as FP8 training, quantization-aware training (QAT), post-training quantization (PTQ), and 2:4 sparsity across different backends.
- Integration with pre-training tools like TorchTitan and fine-tuning platforms like TorchTune and Axolotl ensures a unified workflow from development to deployment in serving environments.
- Facilitated recent launches of quantized Llama 3.21B/3B and LlamaGuard3-8B models by bridging gaps in model optimization pipelines.
- Offers flexibility in targeting diverse hardware environments while enabling efficient model optimization techniques for enhanced performance.
- Open-source project available at https://github.com/pytorch/ao.

Summary1. TorchAO helps make big language models better. 2. It uses PyTorch tricks to help AI models from start to finish. 3. It can improve models in different ways like training with less precision, making them smaller, and using less data. 4. It works with other tools for making and improving models easily. 5. TorchAO is free and can be found online. Definitions- Model optimization: Making models work better or faster. - Framework: A set of tools that help do a job more easily. - Lifecycle: The whole process from the beginning to the end. - Techniques: Different ways of doing something. - Backend: The part of a system that users don't see but helps things work.

TorchAO: PyTorch-Native Training-to-Serving Model Optimization In the rapidly evolving field of artificial intelligence (AI), large language models (LLMs) have gained significant attention for their ability to perform complex natural language processing tasks. However, with the increasing size and complexity of these models, there is a growing need for efficient model optimization techniques that can enhance performance while minimizing resource usage. In their paper titled "TorchAO: PyTorch-Native Training-to-Serving Model Optimization," authors Andrew Or, Apurva Jain, Daniel Vega-Myhre, Jesse Cai, Charles David Hernandez, Zhenrui Zheng, Driss Guessous, Vasiliy Kuznetsov, Christian Puhrsch, Mark Saroufim, Supriya Rao, Thien Tran and Aleksandar Samardžić introduce TorchAO as a comprehensive framework for model optimization in LLMs. The TorchAO framework leverages PyTorch's native support for mixed precision training and quantization techniques to offer an end-to-end workflow for AI models. This allows developers to seamlessly integrate model optimization into the entire lifecycle of LLM development - from pre-training to fine-tuning and deployment in serving environments. One of the key features of TorchAO is its support for various model optimization methods such as FP8 training, quantization-aware training (QAT), post-training quantization (PTQ), and 2:4 sparsity across different hardware backends including server CPU/GPU and mobile/edge devices. This flexibility enables developers to target diverse hardware environments while still being able to efficiently optimize their models. Moreover,TorchAO's integration with pre-training tools like TorchTitan and fine-tuning platforms like TorchTune and Axolotl ensures a unified workflow from model development to deployment in serving environments supported by popular frameworks such as HuggingFace,vLLM,SGLang,and ExecuTorch. This integration bridges gaps in the fragmented space of model optimization pipelines and has already facilitated recent launches of quantized Llama 3.2 1B/3B and LlamaGuard3-8B models. The authors also invite contributions to the open-source project at https://github.com/pytorch/ao, highlighting the collaborative nature of TorchAO's development and its potential for further advancements in model optimization techniques. In conclusion, TorchAO is a comprehensive framework that addresses the growing need for efficient model optimization techniques in LLMs. By leveraging PyTorch's native support for mixed precision training and quantization, it offers an end-to-end workflow for AI models while supporting various hardware environments. Its integration with pre-training tools and fine-tuning platforms ensures a unified workflow from development to deployment, making it a valuable tool for developers working with large language models. With its open-source nature, TorchAO has the potential to drive further advancements in model optimization techniques and contribute to the growth of AI applications across various industries.

Created on 28 Aug. 2025

Assess the quality of the AI-generated content by voting

Score: 0

Similar papers summarized with our AI tools

57.9%

PrefixQuant: Static Quantization Beats Dynamic through Prefixed Outliers in L…

cs.LG

54.7%

QLoRA: Efficient Finetuning of Quantized LLMs

cs.LG

54.4%

FP4 All the Way: Fully Quantized Training of LLMs

cs.LG

52.9%

GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transfor…

cs.LG

51.0%

A Survey on LoRA of Large Language Models

cs.LG

50.6%

Neural Network Quantization for Efficient Inference: A Survey

cs.LG

50.5%

Accuracy is Not All You Need

cs.LG

Navigate through even more similar papers through a

tree representation

Look for similar papers (in beta version)

By clicking on the button above, our algorithm will scan all papers in our database to find the closest based on the contents of the full papers and not just on metadata. Please note that it only works for papers that we have generated summaries for and you can rerun it from time to time to get a more accurate result while our database grows.

Disclaimer: The AI-based summarization tool and virtual assistant provided on this website may not always provide accurate and complete summaries or responses. We encourage you to carefully review and evaluate the generated content to ensure its quality and relevance to your needs.