Optimizing across the entire batch is one secret to better proteins

Optimizing across the entire batch is one secret to better proteins

We break down the ML behind our plate-level design philosophy–how probabilistic modeling, error fingerprints, and portfolio-style optimization help us produce consistently successful experimental rounds from inherently imperfect models.

Nicolas

Nicolas

A key principle that has shaped the development of Cradle is that while we are a software platform, we still produce experiments in the wet lab: modern high-throughput assays operating on batches of 96 or more protein variants. As a result, scientists use Cradle's machine-learning pipeline to generate and evaluate plates of proteins, and we consider it misguided to try and assign value or risk to individual variants.

Editor's note

If you're not familiar with Cradle...

Have a look at our whitepaper, our docs, or our blog, “Unfold”, where you can also subscribe to our newsletter. 


Cradle uses machine learning (ML) to accelerate protein design: We help scientists engineer proteins to become biologics like antibodies or peptides in novel therapeutics; enzymes used in industry for producing more sustainable foods, food ingredients, materials, and pharmaceutical ingredients; and more. To work as such, they need to meet a set of performance specifications. (In therapeutics, such a set is typically referred to as the Target Candidate Profile, or TCP, which is a useful shorthand we'll keep here since enzyme engineering and industrial biocatalysis are also designing toward several properties at once).

We operate in a high-throughput, lab-in-the-loop setting where scientists iteratively assay micro-titer plates of multiples of 96 proteins. Over a few rounds, we'll generate 96, 192, 384 (and so on) protein variants for scientists to populate a plate with, get experimental data on them, feed that data back into the models so they're learning from the scientists' results, and iterate a handful of times until the TCP is reached.

Cradle is more than just a single model that you query for properties like binding affinity; it is a comprehensive system designed to help your protein meet TCP as quickly as possible. We have been very deliberate in defining our optimization objectives as close as possible to the end goal of our users.

In a lead optimization campaign, for instance, our users care about having the best possible protein in as few rounds as possible. So we designed Cradle to take into account the real-life configuration of these rounds and how their success would be evaluated end-to-end.

This article explains what that looks like inside Cradle and how we break down this principle into more precise goals. We'll then explore how this affects how plates are designed. A previously-published companion piece showed what that means for how you should use these designs.

End-to-end lead optimization

Cradle is a software platform, not an ML model. Scientists use our lead-optimization tools to design experimental rounds holistically. A significant part of our technological advantage comes from that extra step of maturity: taking model predictions and transforming them into well-designed experimental rounds.

When a scientist uses Cradle to design a set of sequences, the optimization happens at the set level, not at the individual variant level. This lets us jointly tackle several objectives that cannot be achieved by simple predictions:

Multi-property balance. A plate is designed to maximally cover trade-offs between the objectives you are optimizing for. Across rounds this helps Cradle efficiently learn the landscape and figure out how to reach the desired profile most directly.

Risk management. We’ve observed that the imperfections natural to any model sometimes lead to surprising, beneficial results. Embracing this apparent flaw, we design our models to generate rich, probabilistic predictions that highlight risk profiles–not just expected performance. Greedily picking the best performers, we’ve found, tends to lead to overall structural weaknesses. Conversely, having a firm grip on risk lets us make some chancier moves, with corresponding potentially higher payoff.

These two considerations are typical examples of things that are hard for humans to quantify and balance, but are perfectly suited for ML and probabilistic optimization. Cradle's formalization of these goals, and the automated solution we've built to systematically reach them, are core to the success that scientists using Cradle have had in past campaigns.

Let's look at what this means in more detail.

Multi-property optimization

Deciding how much you’re willing to lose of one thing in exchange for a certain amount of gain of another is a fundamental challenge in any protein engineering campaign. Having invested in user research early on, we know how hard it is to specify preferred trade-offs between different properties upfront. Do you want 10% more activity even if it costs you 15% stability? Would you take 20% more activity in exchange for 20% less stability? Is vastly improved expression worth a modest hit to binding affinity? Sure, you need "good enough" stability and "as high as possible" activity–but what should be the exact exchange rate between them? These trade-offs are context-dependent, campaign-dependent, and often impossible to specify precisely until you have data.

Our constrained optimization setup helps with these decisions. All but one property are defined in terms of a TCP with thresholds to meet. That lets us forecast outcomes in terms of universally comparable trade-offs: probabilities of meeting thresholds. We use a probabilistic modeling approach where we can make statements like, “There is a 30% chance of meeting all constraints, a 20% chance of not being stable enough,” etc. Even a single variant is characterized by a rich probabilistic profile with a very high number (2^N) of possible outcomes.


These N-dimensional representations are a great way to visualize potential experimental outcomes, but they get cumbersome pretty fast.

Constraint probability profile for a single engineered enzyme variant

On our platform, you'll see this rich information displayed in forecasts with dedicated plots.

Good experiment design will balance a plate such that we not only optimize the overall probability of meeting TCP, but we also diversify how we approach TCP. We design plates to scan the Pareto front of trade-offs. For example, some variants might emphasize binding affinity at the cost of some stability, while others might seek balanced improvements across both properties.

This is particularly important because in multi-property optimization, some trade-offs are (at least locally) non-negotiable. You can't always have your cake and eat it too. Maybe you truly can't improve both binding and stability simultaneously from your starting point, at least not without making a larger leap the model isn't confident about yet. But plate-level design helps you understand the shape of what's possible and maximize landscape exploration and opportunities.

In any case, remember that Cradle scans this exponentially large space of outcomes and diversifies over it. That's one of the reasons we don't consider point predictions a very useful output. Cradle provides a much richer representation of how we expect different properties to interact.


If you cherry-pick variants based on a single property or a simplistic scoring function, you collapse this rich data down to a single point in property space. That approach squanders opportunities to learn about the landscape and make informed decisions about which compromises are acceptable.

Risk management

As mentioned above, a key component of our modeling approach is a probabilistic understanding of predictions. In short, this means that we don't predict a single value for variants, but rather work to anticipate modeling errors and to quantify how big these will be.

Selecting for mistakes might seem an odd way of doing things for a company building machine learning tools for protein engineering, but it's actually essential to delivering real value. Models will make errors, and they will be wrong in systematic ways. A surprisingly large part of our job is to design experimental strategies that remain robust and informative even when the models aren't perfect–we want to make the most out of something we know is sometimes going to be wrong. In other words, Cradle is fundamentally a tool to extract maximum utility out of inherently imperfect models.

That doesn't mean that our models aren’t state-of-the-art–they absolutely are. And we are continuously improving them, both through ML R&D and data acquisition in our wet lab. Nevertheless, the adage "all models are wrong, some are useful" needs to be taken seriously–especially in biology, where every rule has an exception, everything is complicated and depends on everything else, and not everything is understood well. Until protein biochemical and functional data increases drastically both in depth and in breadth, predictive models will have weaknesses. We need to account for those weaknesses, or we risk running afoul of them.

Having a mature and broadly-applicable lead-optimization toolbox means we’re able to tackle both bread-and-butter predictions like melting temperature as well as more niche assays, such as functional readouts that are relevant to a unique protein family in a rare host. As a result, Cradle is designed to handle models that range from very accurate to imprecise, and dynamically manages risks and opportunities when designing your next round.

For example, if the model is more confident than ends up being justified about a particular structural motif or substitution pattern, we don't want all 96 variants betting on that same assumption. We need to instead evaluate a diverse panel of variants, in order to manage modeling risk. This is one aspect of plate design that's usually easy to communicate: We get more shots on goal by spreading the ball around.

Traditionally, this diversification is assessed at the sequence level by managing overlaps between mutations. You might use heuristics like "ensure variants differ by at least N mutations" or "sample from different clusters in sequence space" so that your 96 variants don’t share the same five critical substitutions–because if those substitutions don't work as expected, your entire plate fails.

A sensible approach, but quite limited. Not all mutations are equal in terms of either their impact or the model's uncertainty about them. More fundamentally, this type of sequence-level diversification is using a heuristic (sequence diversity) as a proxy for an end-goal (hedging modeling risk). Sequence diversity is easy to calculate, but it's not quite what we care about.

Cradle’s fine-grained description of expected modeling errors lets us do something more sophisticated. We can characterize similarity between variants in terms of "error similarity," i.e., whether the model is likely to make similar mistakes on them. We can make statements like: "If my prediction is over-optimistic for variant 12, it's likely I'm also over-optimistic for variant 15, even though these variants differ by 5 mutations."

In broad strokes, we predict such "error fingerprints" for each variant. These capture which aspects of the model's internals are driving the prediction, which structural assumptions are being relied upon, and where the uncertainty is coming from. We can then use the fingerprints to analyze joint over- or under-optimism across the plate.

This lets us look directly at something analogous to "structural risks" in financial portfolio management. If your investment advisor overweights in consumer goods because the economy looks healthy, you’re in trouble when inflation or unemployment numbers disappoint. Similarly, if an entire plate depends on certain implicit modeling choices, then "structural over-optimism" can lead to experimental results disappointing across the board. By selecting groups of sequences where we expect errors to be anti-correlated, we can leverage "error fingerprints" to not only diversify, but also to hedge risks. This lets us optimize directly for balanced risk- and opportunity-taking.

In our in-house experiments–many rounds over years while continuously improving the model–we’ve found that this approach still leads to sequence diversity. Simply because variants with very different error fingerprints usually also have rather different sequences. But the diversification rules are more sophisticated and harder to characterize with simple sequence-based heuristics. Interestingly, we've found that these fuzzy, fingerprint-based diversification rules are also much harder for optimization algorithms to game. A simple sequence diversity constraint can sometimes be satisfied by superficial changes that don't actually change the modeling risk profile; error-fingerprint diversification cuts through to what actually matters.

Crucially, "risk-balanced" is a property of the whole plate. A particular variant might incur specific modeling risks, but that's acceptable if other variants in the plate compensate for those risks by betting on different assumptions. One variant might be a high-risk bet on an under-explored region of sequence space, but it's hedged by more conservative variants that the model is confident about for different reasons.

This is why our software designs experiments at the plate level: The plate is constructed so that the model's blind spots are covered by variants that depend on different assumptions. Collectively, a diverse basket of designs compensate for each other's weaknesses. Just like Treasury bonds compensate for tech stocks.

Conclusion

Multi-property balance and risk-balanced design are two sides of the same coin: The right unit of optimization in a lead campaign isn't a sequence, it's a plate. Since proteins move through the lab in batches, once you accept that models are imperfect by nature, it becomes clear that single-variant thinking is a misalignment between what you're optimizing and what you actually care about. The best individual prediction won't reliably get you to a target product profile faster; a well-constructed portfolio almost always will.

This is why Cradle steers you toward plate-level forecasts. Ranks can lead to the collapse of a rich, multi-dimensional view of the design landscape into a single point in property space. Information essential to learning from the round is lost, and you concentrate experimental risk on whichever modeling assumptions the top-ranked candidates happen to share. These costs compound across rounds.

Trusting plate-level design means accepting that an individual variant might not be the one you'd have picked on its own. But over a campaign, the plate moves you toward your TCP—and it's the experimental unit we've spent years learning to optimize.

A key principle that has shaped the development of Cradle is that while we are a software platform, we still produce experiments in the wet lab: modern high-throughput assays operating on batches of 96 or more protein variants. As a result, scientists use Cradle's machine-learning pipeline to generate and evaluate plates of proteins, and we consider it misguided to try and assign value or risk to individual variants.

Editor's note

If you're not familiar with Cradle...

Have a look at our whitepaper, our docs, or our blog, “Unfold”, where you can also subscribe to our newsletter. 


Cradle uses machine learning (ML) to accelerate protein design: We help scientists engineer proteins to become biologics like antibodies or peptides in novel therapeutics; enzymes used in industry for producing more sustainable foods, food ingredients, materials, and pharmaceutical ingredients; and more. To work as such, they need to meet a set of performance specifications. (In therapeutics, such a set is typically referred to as the Target Candidate Profile, or TCP, which is a useful shorthand we'll keep here since enzyme engineering and industrial biocatalysis are also designing toward several properties at once).

We operate in a high-throughput, lab-in-the-loop setting where scientists iteratively assay micro-titer plates of multiples of 96 proteins. Over a few rounds, we'll generate 96, 192, 384 (and so on) protein variants for scientists to populate a plate with, get experimental data on them, feed that data back into the models so they're learning from the scientists' results, and iterate a handful of times until the TCP is reached.

Cradle is more than just a single model that you query for properties like binding affinity; it is a comprehensive system designed to help your protein meet TCP as quickly as possible. We have been very deliberate in defining our optimization objectives as close as possible to the end goal of our users.

In a lead optimization campaign, for instance, our users care about having the best possible protein in as few rounds as possible. So we designed Cradle to take into account the real-life configuration of these rounds and how their success would be evaluated end-to-end.

This article explains what that looks like inside Cradle and how we break down this principle into more precise goals. We'll then explore how this affects how plates are designed. A previously-published companion piece showed what that means for how you should use these designs.

End-to-end lead optimization

Cradle is a software platform, not an ML model. Scientists use our lead-optimization tools to design experimental rounds holistically. A significant part of our technological advantage comes from that extra step of maturity: taking model predictions and transforming them into well-designed experimental rounds.

When a scientist uses Cradle to design a set of sequences, the optimization happens at the set level, not at the individual variant level. This lets us jointly tackle several objectives that cannot be achieved by simple predictions:

Multi-property balance. A plate is designed to maximally cover trade-offs between the objectives you are optimizing for. Across rounds this helps Cradle efficiently learn the landscape and figure out how to reach the desired profile most directly.

Risk management. We’ve observed that the imperfections natural to any model sometimes lead to surprising, beneficial results. Embracing this apparent flaw, we design our models to generate rich, probabilistic predictions that highlight risk profiles–not just expected performance. Greedily picking the best performers, we’ve found, tends to lead to overall structural weaknesses. Conversely, having a firm grip on risk lets us make some chancier moves, with corresponding potentially higher payoff.

These two considerations are typical examples of things that are hard for humans to quantify and balance, but are perfectly suited for ML and probabilistic optimization. Cradle's formalization of these goals, and the automated solution we've built to systematically reach them, are core to the success that scientists using Cradle have had in past campaigns.

Let's look at what this means in more detail.

Multi-property optimization

Deciding how much you’re willing to lose of one thing in exchange for a certain amount of gain of another is a fundamental challenge in any protein engineering campaign. Having invested in user research early on, we know how hard it is to specify preferred trade-offs between different properties upfront. Do you want 10% more activity even if it costs you 15% stability? Would you take 20% more activity in exchange for 20% less stability? Is vastly improved expression worth a modest hit to binding affinity? Sure, you need "good enough" stability and "as high as possible" activity–but what should be the exact exchange rate between them? These trade-offs are context-dependent, campaign-dependent, and often impossible to specify precisely until you have data.

Our constrained optimization setup helps with these decisions. All but one property are defined in terms of a TCP with thresholds to meet. That lets us forecast outcomes in terms of universally comparable trade-offs: probabilities of meeting thresholds. We use a probabilistic modeling approach where we can make statements like, “There is a 30% chance of meeting all constraints, a 20% chance of not being stable enough,” etc. Even a single variant is characterized by a rich probabilistic profile with a very high number (2^N) of possible outcomes.


These N-dimensional representations are a great way to visualize potential experimental outcomes, but they get cumbersome pretty fast.

Constraint probability profile for a single engineered enzyme variant

On our platform, you'll see this rich information displayed in forecasts with dedicated plots.

Good experiment design will balance a plate such that we not only optimize the overall probability of meeting TCP, but we also diversify how we approach TCP. We design plates to scan the Pareto front of trade-offs. For example, some variants might emphasize binding affinity at the cost of some stability, while others might seek balanced improvements across both properties.

This is particularly important because in multi-property optimization, some trade-offs are (at least locally) non-negotiable. You can't always have your cake and eat it too. Maybe you truly can't improve both binding and stability simultaneously from your starting point, at least not without making a larger leap the model isn't confident about yet. But plate-level design helps you understand the shape of what's possible and maximize landscape exploration and opportunities.

In any case, remember that Cradle scans this exponentially large space of outcomes and diversifies over it. That's one of the reasons we don't consider point predictions a very useful output. Cradle provides a much richer representation of how we expect different properties to interact.


If you cherry-pick variants based on a single property or a simplistic scoring function, you collapse this rich data down to a single point in property space. That approach squanders opportunities to learn about the landscape and make informed decisions about which compromises are acceptable.

Risk management

As mentioned above, a key component of our modeling approach is a probabilistic understanding of predictions. In short, this means that we don't predict a single value for variants, but rather work to anticipate modeling errors and to quantify how big these will be.

Selecting for mistakes might seem an odd way of doing things for a company building machine learning tools for protein engineering, but it's actually essential to delivering real value. Models will make errors, and they will be wrong in systematic ways. A surprisingly large part of our job is to design experimental strategies that remain robust and informative even when the models aren't perfect–we want to make the most out of something we know is sometimes going to be wrong. In other words, Cradle is fundamentally a tool to extract maximum utility out of inherently imperfect models.

That doesn't mean that our models aren’t state-of-the-art–they absolutely are. And we are continuously improving them, both through ML R&D and data acquisition in our wet lab. Nevertheless, the adage "all models are wrong, some are useful" needs to be taken seriously–especially in biology, where every rule has an exception, everything is complicated and depends on everything else, and not everything is understood well. Until protein biochemical and functional data increases drastically both in depth and in breadth, predictive models will have weaknesses. We need to account for those weaknesses, or we risk running afoul of them.

Having a mature and broadly-applicable lead-optimization toolbox means we’re able to tackle both bread-and-butter predictions like melting temperature as well as more niche assays, such as functional readouts that are relevant to a unique protein family in a rare host. As a result, Cradle is designed to handle models that range from very accurate to imprecise, and dynamically manages risks and opportunities when designing your next round.

For example, if the model is more confident than ends up being justified about a particular structural motif or substitution pattern, we don't want all 96 variants betting on that same assumption. We need to instead evaluate a diverse panel of variants, in order to manage modeling risk. This is one aspect of plate design that's usually easy to communicate: We get more shots on goal by spreading the ball around.

Traditionally, this diversification is assessed at the sequence level by managing overlaps between mutations. You might use heuristics like "ensure variants differ by at least N mutations" or "sample from different clusters in sequence space" so that your 96 variants don’t share the same five critical substitutions–because if those substitutions don't work as expected, your entire plate fails.

A sensible approach, but quite limited. Not all mutations are equal in terms of either their impact or the model's uncertainty about them. More fundamentally, this type of sequence-level diversification is using a heuristic (sequence diversity) as a proxy for an end-goal (hedging modeling risk). Sequence diversity is easy to calculate, but it's not quite what we care about.

Cradle’s fine-grained description of expected modeling errors lets us do something more sophisticated. We can characterize similarity between variants in terms of "error similarity," i.e., whether the model is likely to make similar mistakes on them. We can make statements like: "If my prediction is over-optimistic for variant 12, it's likely I'm also over-optimistic for variant 15, even though these variants differ by 5 mutations."

In broad strokes, we predict such "error fingerprints" for each variant. These capture which aspects of the model's internals are driving the prediction, which structural assumptions are being relied upon, and where the uncertainty is coming from. We can then use the fingerprints to analyze joint over- or under-optimism across the plate.

This lets us look directly at something analogous to "structural risks" in financial portfolio management. If your investment advisor overweights in consumer goods because the economy looks healthy, you’re in trouble when inflation or unemployment numbers disappoint. Similarly, if an entire plate depends on certain implicit modeling choices, then "structural over-optimism" can lead to experimental results disappointing across the board. By selecting groups of sequences where we expect errors to be anti-correlated, we can leverage "error fingerprints" to not only diversify, but also to hedge risks. This lets us optimize directly for balanced risk- and opportunity-taking.

In our in-house experiments–many rounds over years while continuously improving the model–we’ve found that this approach still leads to sequence diversity. Simply because variants with very different error fingerprints usually also have rather different sequences. But the diversification rules are more sophisticated and harder to characterize with simple sequence-based heuristics. Interestingly, we've found that these fuzzy, fingerprint-based diversification rules are also much harder for optimization algorithms to game. A simple sequence diversity constraint can sometimes be satisfied by superficial changes that don't actually change the modeling risk profile; error-fingerprint diversification cuts through to what actually matters.

Crucially, "risk-balanced" is a property of the whole plate. A particular variant might incur specific modeling risks, but that's acceptable if other variants in the plate compensate for those risks by betting on different assumptions. One variant might be a high-risk bet on an under-explored region of sequence space, but it's hedged by more conservative variants that the model is confident about for different reasons.

This is why our software designs experiments at the plate level: The plate is constructed so that the model's blind spots are covered by variants that depend on different assumptions. Collectively, a diverse basket of designs compensate for each other's weaknesses. Just like Treasury bonds compensate for tech stocks.

Conclusion

Multi-property balance and risk-balanced design are two sides of the same coin: The right unit of optimization in a lead campaign isn't a sequence, it's a plate. Since proteins move through the lab in batches, once you accept that models are imperfect by nature, it becomes clear that single-variant thinking is a misalignment between what you're optimizing and what you actually care about. The best individual prediction won't reliably get you to a target product profile faster; a well-constructed portfolio almost always will.

This is why Cradle steers you toward plate-level forecasts. Ranks can lead to the collapse of a rich, multi-dimensional view of the design landscape into a single point in property space. Information essential to learning from the round is lost, and you concentrate experimental risk on whichever modeling assumptions the top-ranked candidates happen to share. These costs compound across rounds.

Trusting plate-level design means accepting that an individual variant might not be the one you'd have picked on its own. But over a campaign, the plate moves you toward your TCP—and it's the experimental unit we've spent years learning to optimize.