Skip to main content

What Are Scaling Laws in AI?

What Are Scaling Laws in AI?

Scaling laws in artificial intelligence (AI) are different rules that describe the impact on performance and resource requirements of AI systems as their architecture scales. In statistics, a power law is a functional relationship between two data points, where a change in one data point results in a relative and proportional change in the other data point—one quantity varies as a mathematical power function of another. Scaling laws in AI attempt to identify and establish such predictable relationships within the context of AI. They take AI factors like model parameters, training data, computational resources, etc., and establish mathematical relationships between them. Neural scaling law is another term for applying scaling laws in artificial intelligence.

Why are scaling laws important in AI research?

Neural scaling laws are a new field of research, and AI companies publish papers and identify relationships after applying rigorous empirical methods to their deep learning systems. The scaling laws bring predictability to AI research and development. Data scientists can pre-calculate project costs and estimate performance to some extent. They can also test and see if any planned efforts will bring the returns they expect. We explain the benefits below.

Lower risk in AI research

Training larger models with millions and billions of parameters is expensive and can cost millions of dollars. Scaling laws help researchers better estimate training data and model size, required computational resources, and timelines and potential costs. They can determine in advance if the expenditure will bring the required gains or not.

Select the best models

Many organizations prefer using existing foundation models instead of training new ones from scratch. Given the range and variety of models available, scaling laws help determine the best choices for your specific task. For instance, smaller models offer more deployment options, are less expensive, and process data faster, but may be less accurate. The mathematical foundation of scaling laws can help AI teams make better decisions.

Optimize model training

Teams use the mathematical equations in scaling laws to create an optimal training configuration for their large language models. For example, in certain scenarios, even large models may require a relatively small dataset for optimum performance. There is a certain threshold beyond which more training effort does not give the required returns to make it worthwhile. Scaling laws help identify such thresholds and maximize efficiency in the training process.

What are the main parameters in scaling laws?

A neural scaling law relates the parameters of a family of neural networks. In general, you can characterize a neural network model (or algorithm) by four parameters:

  • Model size
  • Training dataset size
  • Performance after training
  • Cost of training

You can precisely define each of these four variables into a real number. The neural scaling laws empirically establish that the four parameters are related by statistical rules, describable as mathematical equations. In explaining neural scaling laws, it is important to understand the four key parameters.

Model size

Model size refers to how many parameters the model has. A parameter represents the neural connections the model makes during the training process. In the context of generative AI and foundation models (FMs), parameters are the trainable weights and biases within the model. They represent the model's ability to perform natural language processing tasks. For example, parameters measure a large language model's ability to generate human language or a generative image model's ability to create accurate images from text prompts.

Training dataset size

Training dataset size is quantified by the number of data points used for training the model. Initial or pre-training datasets are important, but most modern organizations are moving towards adapting foundation models for custom applications that use internal data. This custom or fine-tuning dataset size is more important so organizations can understand and prepare the data they need for their AI applications. It has been found that a small amount of high-quality data suffices for fine-tuning, and more data does not improve performance.

Performance after training

You can evaluate the performance of neural language models based on their ability to accurately predict or generate output for a given unknown input. Common model performance metrics include accuracy, precision, Elo rating in a competition against other models, or preference by a human judge. Performance is improved by increasing training data size, model size, and other algorithmic optimization methods.

Training cost

The training cost is measured in the time and computational resources required for model training. It is impacted by several factors, including model size, training data size, model complexity, training data complexity, and more. However, the relationship between the training dataset and training cost is not linear. For example, doubling the training data may not necessarily double the training cost because of how AI models train.

What are some guidelines established by scaling laws?

The empirical findings and published AI research present guidelines to better direct AI development and achieve artificial general intelligence. We present some findings below. However, the field is evolving rapidly, and things can change at any time.

Performance and model size

As the number of parameters in a language model increases, its performance on a wide range of tasks generally improves. This includes a better understanding of context, more accurate text generation, and an improved ability to learn from fewer examples. However, there is a phenomenon of diminishing returns. This means that each additional parameter added to the model contributes less to its overall performance improvement than the previous one.

For instance, in early 2020, research organizations focused on model size, building language models like GPT-3 and BLOOM with around 175 billion parameters, while MT-NLG was trained on 530 billion parameters. However, in 2022, it was observed that the current balance of compute between model parameters and training dataset size was suboptimal. Newer empirical scaling laws suggest that smaller models trained on more data had the potential to outperform large models. For example, Chinchilla, with 70B parameters, is known to outperform much bigger models.

Training data and model size

Larger models require significantly more data to train effectively. They also require exponentially more computational resources, making them increasingly expensive and energy-intensive. Training very large models also introduces new challenges in optimization. It becomes increasingly difficult to find the best settings and techniques to efficiently train these models without running into issues like overfitting or training instabilities. Having said that, it is important to note that larger models are more sample-efficient and require less training time to reach better performance.

Performance and model architecture

Model architecture refers to the specific design and structure of the model— such as the choice of neurons in the neural networks, how they are organized and connected (number of layers, connections between layers), and other structural aspects of the model. Deep-learning scaling laws observed that model architecture plays a weak role in model performance. Model performance depends most strongly on the number of model parameters, the size of the training dataset, and the amount of compute used for training.

The only exceptions are scenarios when the architecture itself becomes a bottleneck. Yet, within reasonable limits, performance depends very weakly on other architectural hyperparameters, such as the number of model layers and the number of neurons in each layer.

Selecting the right model

When fine-tuning foundation models for specific use cases, it is important to note that the models take up physical space. For example, smaller models run on a single GPU, but larger models may run on two or more GPUs. As accelerators and GPUs go up, accuracy improves, but the inference or run time of the model and its operational costs also increase. In contrast, smaller models give you more deployment options, are less expensive, and offer faster inference time. However, they may have reduced output accuracy. The trade-off between accuracy and cost/time needs to be resolved by AI teams based on business requirements.

Diagram of model selection trade-offs across speed, precision, and cost for foundation models FM1, FM2, and FM3, showing that higher speed favors smaller lower-cost models while higher precision favors larger higher-cost models, with FM2 selected.

What are some challenges in establishing scaling laws?

While scaling laws offer guidelines, they also present some contradictions.

Linear scaling laws

The scaling laws seem to present mathematical linearity with respect to model size, data size, and compute budget. This is contradictory because language, the fundamental domain of these models, has a degree of unpredictability or complexity. As models scale up, you expect them to approach a limit where they capture most of this complexity but not all. However, if the scaling laws are linear, the models continue to improve without approaching the theoretical limit of language understanding. This seems counterintuitive, given the inherent complexity and unpredictability of language.

Overfitting

Overfitting is when a model learns the training data too well, including its noise and outliers, which harms its ability to generalize to new data. For example, a model trained only on cat pictures will not be able to identify dogs. Scaling law research showed that even when a model is trained with just one complete pass through the training data, without reusing data, it can still overfit. This finding is somewhat surprising because, at a smaller scale, overfitting is more likely with extensive training and data re-usage.

Inconsistency in compute efficiency

The scaling laws are inconsistent in the relationship between the compute required for training and model performance. Certain mathematical exponents in the scaling equation vary significantly, indicating that the efficiency of compute usage in improving model performance is not consistent and can change dramatically under different conditions or constraints. This contradicts the idea of a stable, predictable scaling law for compute efficiency.

How can AWS help with your AI efforts?

AWS can help you at every stage of your AI adoption journey with the most comprehensive set of artificial intelligence (AI) services, infrastructure, and implementation resources. Enhance customer experiences, enable faster and better decision-making, and optimize business processes with AI on AWS. For example, you can use:

  • Amazon Bedrock to select, customize, train, and deploy industry-leading foundational models with proprietary data.
  • Amazon CodeWhisperer to generate code suggestions—ranging from snippets to full functions—in real-time in the IDE based on your comments and existing code.
  • Amazon Q to get fast, relevant answers to pressing questions, solve problems and generate content. You can also take action using the data and expertise found in your company's information repositories, code, and enterprise systems.
  • Amazon SageMaker Jumpstart to accelerate AI development by building, training, and deploying foundational models in a machine-learning hub.

Get started with AI on AWS by creating a free account today.

Browse all cloud computing concepts

Browse all cloud computing concepts content here:

Loading
Loading
Loading
Loading
Loading

Did you find what you were looking for today?

Let us know so we can improve the quality of the content on our pages