10 Concepts You Should Know Regarding ML in Production | Deepchecks

10 Concepts You Should Know Regarding ML in Production

Itay Gabbay

|
April 07, 2021 | 7 mins

Introduction

Deploying Machine Learning (ML) models to production poses a variety of challenges that may be related to different components of the system. This article aims to help you navigate through the world of ML in production by providing a basic understanding of the concepts that will give you control of your ML models even after they have been deployed.

1. Observability and Monitoring

To be in control of your ML model in production, it is essential to receive live information regarding performance. This is done by setting up a dashboard with live information regarding model evaluation metrics (i.e., uptime, resource utilization, latency, and notifications) that pop up using customized triggers.

Observability requires that part of the system can be observed in action – the input data, engineered features, and finally, model performance and predictions. When it comes to observability, the rule is “the more the merrier.”

2. Data Integrity Issues

Also known as Training-serving Skew and Data Skew.

In most ML applications, there is a long and complex process that generates and transforms the data that ends up being fed to our model. Some of these stages may not even be under our control, so a new deployment in another company’s website can greatly influence the final format of the input to our model. Even a minor change, like a field rename or an introduction of a new value to the gender field, may cause our model to perform poorly, and we may never be notified of this issue.

Data integrity issues can potentially destroy our model, but simple steps can be taken to eliminate that threat.

3. Model Degradation and Staleness

As Ecclesiastes said, “To everything, there is a season and a time to every purpose under the heavens, so too for your ML model!”

An ML model is only as good as the data it’s trained on. As the world changes and new data shifts, your model will likely go stale and degrade. This process is perfectly normal and can take varying lengths of time, depending on the scenario.

These processes can be detected by directly measuring the degradation in performance metrics, by estimating the length of the process based on historical data, or by detecting data drift.

4. Data Drift and Concept Drift

Data drift and concept drift are the most common causes of model degradation.

Data Drift

When P(X). The distribution of features changes over time. This can happen either because of some shift in the data structure, or because of a change in the real world. For example, the average profile for a person requesting a loan might change in light of a financial crisis.

Data drift can be detected in real-time, even when the real labels are not available, so it serves as a signal for model degradation in any scenario.

Concept Drift

When P(Y|X). The distribution of correct labels given the features changes over time. This, too, can be caused by a shift in the data structure or by a change in reality but affects prediction quality indefinitely.

5. Feedback Loop

Ever heard that horrible noise when a mic is too close to the speaker? That can happen to your ML model as well.

6. Model Retraining

The availability of new labeled data and the degradation of our production model indicate it is time to retrain the model. The model is typically trained from scratch on the full dataset, however, there are paradigms such as incremental learning and online learning that attempt to update an existing model by training on new examples as they become available instead of training the model from scratch.

7. Seasonality and Data Fluctuation Patterns

Detecting concept drift or decrease in model performance is not the end of the game. We must ask ourselves what might have caused the shift to understand whether our model will “go back to normal,” and whether we should create a more robust model that won’t undergo the same degradation process.

8. Batch Processing vs. Real-time Processing

Batch Processing

Feeding the model batches of samples of a set size is typically more efficient since we can optimize the parallelization capabilities by selecting the best batch size. However, this is not a good option if we are supposed to process the request and make a prediction in real-time.

Real-time Processing or Stream Processing

In a typical setting, we receive a request from a user, preprocess the request, feed it to the ML model, post-process, then return the result in real-time. Latency can have a significant negative impact, so you may need to compress your model and try to minimize operations per request to keep the latency to a minimum.

9. Model Compression

Model Compression enables quicker predictions, reduces latency, and reduces the memory footprint of the model where the memory or disk size is limited (e.g., on-device ML). This can be done in multiple ways:

Quantization

This reduces the floating-point precision for each parameter.

Pruning

Neural Network Pruning is the process of eliminating parameters of the network iteratively to compress the model and improve inference speed without hurting accuracy.

Knowledge Distillation

Using a teacher-student model, we can create an equivalent model that is simpler and smaller (the student) that learns to imitate the larger model (teacher).

10. A/B Testing

A/B Testing compares two variants of a product to test response to a specific feature. In a similar fashion, we can use A/B testing for our ML models, thus enabling evaluation of whether the newer model actually achieves better performance in scenarios where we don’t have full information.

Conclusion

These basic concepts regarding ML systems in production will hopefully allow you to start navigating this sea and take control of your models before it’s too late.