Data Predictability Shapes Weibull Weight-Scale Growth in Transformer Training

AI-generated keywords: Transformer weight dynamics Weibull distribution Learning-rate-conditioned law Pre-training data predictability Neural network models

AI-generated Key Points

  • Two-parameter Weibull distribution effectively summarizes weight magnitudes
  • Scale parameter λ crucial for capturing training-induced changes in weight magnitudes
  • Learning-rate-conditioned law governs how scale parameter λ grows: λ^2 - λ0^2 = C0(η) + C1(η)(Hr - D)^0.59
  • Forward predictive model successfully recovers held-out within-family weight growth with 5.7% relative error
  • Importance of understanding classical metrics and estimator saturation in training dynamics
  • Correlations along the corruption axis do not imply causal mediation, but indicate predictive relations between pre-training data statistics and post-training weight-scale outcomes
Also access our AI generated: Comprehensive summary, Lay summary, Blog-like article; or ask questions about this paper to our AI assistant.

Authors: Tiexin Ding

27 pages, 14 figures, 5 tables. Code and data: https://github.com/tiexinding/NPM-Weibull-public
License: CC BY 4.0

Abstract: A trained transformer's weight magnitudes can be summarized by a two-parameter Weibull distribution whose shape $k \approx 1.2$ is stable across layers and models, so the scale $λ$ carries most training-induced movement. What corpus property sets how much $λ$ grows? Using the bigram conditional entropy $D = H(\text{next} \mid \text{prev})$, a training-free statistic computed before training, we find across controlled corruption families a learning-rate-conditioned law, $λ^2 - λ_0^2 = C_0(η) + C_1(η)(H_r - D)^{0.59}$, where $H_r$ is a matched-budget shuffle baseline. The convex exponent is inherited from an independently measured data-side saturation relation rather than fitted directly to the growth curve. After removing the two per-$η$ coefficients, 23 runs spanning an order of magnitude in learning rate collapse onto $(H_r - D)^{0.59}$ with unit slope ($R^2 = 0.941$; direct per-$η$ fits are weaker, $R^2 \approx 0.82$). Because $D$ is computed before training, the law is a forward predictor: an end-to-end self-validation recovers held-out within-family weight growth with 5.7% relative error. The readout holds at model and per-layer resolutions and across two tested architectures, with the functional form preserved and only the coefficients changing. It also marks its boundary: cross-corpus prediction over-predicts code, implicating redundancy as a second axis of a broader $Φ(D,R,A,H)$ data-to-weight framework.

Submitted to arXiv on 27 Jun. 2026

Ask questions about this paper to our AI assistant

You can also chat with multiple papers at once here.

AI assistant instructions?

Results of the summarizing process for the arXiv paper: 2608.23573v1

In this study, the researchers delve into the dynamics of transformer weight magnitudes and their relationship to training-induced movement. They discover that a two-parameter Weibull distribution effectively summarizes the weight magnitudes, with a stable shape parameter $k \approx 1.2$ across layers and models. The scale parameter $\lambda$ plays a crucial role in capturing training-induced changes in weight magnitudes. By examining the bigram conditional entropy $D = H(\text{next} \mid \text{prev})$ before training, the researchers uncover a learning-rate-conditioned law that governs how the scale parameter $\lambda$ grows. This law is expressed as $\lambda^2 - \lambda_0^2 = C_0(\eta) + C_1(\eta)(H_r - D)^{0.59}$, where $H_r$ represents a matched-budget shuffle baseline. The convex exponent of 0.59 is inherited from an independently measured data-side saturation relation, providing insight into the growth curve without direct fitting. Further analysis reveals that after removing per-learning-rate coefficients, multiple runs with varying learning rates collapse onto $(H_r - D)^{0.59}$ with high accuracy ($R^2 = 0.941$). This forward predictive model successfully recovers held-out within-family weight growth with only a 5.7% relative error, showcasing its effectiveness as a predictor. The study also explores classical metrics and estimator saturation, highlighting the importance of understanding how these metrics interact with training dynamics. Replication and causal scope are discussed to emphasize that correlations along the corruption axis do not imply causal mediation but rather indicate predictive relations between pre-training data statistics and post-training weight-scale outcomes. In conclusion, this research establishes a learning-rate-conditioned law linking pre-training data predictability to Weibull weight-scale parameter growth. By bridging measurable corpus statistics with weight-scale dynamics, this study sets the foundation for a broader framework (Φ(D, R, A, H)) where data structure, redundancy, architecture, and hyperparameters can be studied systematically one dimension at a time. The findings presented in this study provide valuable insights into understanding how pre-training data properties influence weight dynamics in transformers and offer a structured approach for analyzing these complex interactions in neural network models.
Created on 31 Aug. 2026

Assess the quality of the AI-generated content by voting

Score: 0

Why do we need votes?

Votes are used to determine whether we need to re-run our summarizing tools. If the count reaches -10, our tools can be restarted.

Similar papers summarized with our AI tools

Navigate through even more similar papers through a

tree representation

Look for similar papers (in beta version)

By clicking on the button above, our algorithm will scan all papers in our database to find the closest based on the contents of the full papers and not just on metadata. Please note that it only works for papers that we have generated summaries for and you can rerun it from time to time to get a more accurate result while our database grows.

Disclaimer: The AI-based summarization tool and virtual assistant provided on this website may not always provide accurate and complete summaries or responses. We encourage you to carefully review and evaluate the generated content to ensure its quality and relevance to your needs.