We've talked with researchers at early-stage materials companies and academic-adjacent labs who all describe some version of the same situation: a quarterly synthesis budget, a candidate list longer than the budget can cover, and a process for deciding which candidates get made that amounts to: the senior researcher's intuition plus what's been in the literature recently.
The hit rate, meaning the fraction of synthesized candidates that meet the target property spec, is typically 20-35% in the early phases of a discovery campaign. Sometimes worse. Which means 65-80% of synthesis work, of chemist time, of characterization instrument time, of consumable cost, goes toward confirming that something doesn't work.
This is presented as unavoidable. The argument is: discovery is hard, you have to fail to learn, the hit rate will improve as you accumulate knowledge. All of this is true. But the unstated assumption is that the failure is evenly distributed and that every negative result is informative. It's often not. Many synthesis failures are predictable before the experiment if the right information were available and organized.
The information problem, not a resources problem
The 70% failure rate is not primarily because synthesis is expensive or because labs are understaffed. It's because the information needed to distinguish high-probability candidates from low-probability ones exists but is not accessible in the form that's useful at the decision point.
That information is: the relationship between composition and target property across the space you're exploring. Some of it exists in the scientific literature, scattered across papers with inconsistent measurement protocols. Some exists in DFT databases as computed properties that correlate imperfectly with experimental ones. Some exists in your own lab's historical data, often partially curated, sometimes in lab notebooks that haven't been digitized.
The practical problem is not that this information doesn't exist. It's that assembling it into a form that ranks candidates by expected performance, with uncertainty bounds that tell you how confident to be, is a data integration and modeling task that most R&D teams aren't set up to do as a routine part of experiment planning.
That's the gap we're trying to close at Matforgelab. Not by replacing experimental intuition, but by making the relevant information available before the synthesis queue is committed.
What "front-loading" actually means in practice
Front-loading means shifting prediction work to before synthesis rather than after. Before a candidate enters the synthesis queue, it gets a property prediction with uncertainty bounds. Candidates with predicted properties outside the target window, or with predicted properties in-window but very high uncertainty, are handled differently than candidates with confident in-window predictions.
Specifically, a typical front-loaded screening workflow looks like this. A candidate set is generated, either by systematic substitution on a known parent composition, by literature mining for candidates in adjacent chemistry space, or by a generative model. Each candidate is evaluated by a property prediction model. Candidates fall into three buckets:
Bucket one: high predicted performance, low uncertainty. These go into the synthesis queue with high priority. The prediction is confident that they meet the target spec. Synthesis is likely to confirm.
Bucket two: predicted performance near the threshold, moderate uncertainty. These are information-valuable candidates. Synthesizing them will sharpen the model's boundary between passing and failing compositions. They go into the queue at medium priority, explicitly framed as "model-updating experiments" rather than "we expect these to work."
Bucket three: low predicted performance, regardless of uncertainty. These get cut. The prediction may be wrong for any individual candidate, but across the set, eliminating the low-predicted-performance tail substantially improves campaign hit rates. This is where most of the savings come from.
An honest accounting of what prediction can and cannot eliminate
Prediction-based front-loading doesn't eliminate synthesis failures. It reduces their fraction and makes them more informative when they happen.
A composition model can confidently rule out candidates with thermodynamic instabilities: compositions that will phase-separate rather than form the target phase, or that have known competing phases at the synthesis temperature. These are hard eliminates based on well-understood thermodynamics, and ruling them out before synthesis is straightforward with the right model.
A composition model is less reliable at predicting properties dominated by processing-sensitive variables. As noted in our earlier post on cycle-life prediction, properties like grain boundary conductance, fracture toughness, and thin-film morphology depend on synthesis and processing conditions in ways that bulk composition doesn't capture. For these, prediction can narrow the field but not confidently rank within the remaining candidates.
We're also not saying that high predicted performance means the synthesis will succeed. Synthesis can fail for reasons unrelated to intrinsic material properties: impurity phases from contamination, failed sintering due to equipment variability, measurement artifacts. Prediction screens out the intrinsically bad candidates; it doesn't fix lab execution problems.
The cumulative cost argument
The economic argument for prediction-based screening is clearest when you look at it cumulatively across a discovery program rather than per-experiment.
Consider a 12-month discovery program with a monthly synthesis budget of 20 samples. At 25% hit rate without prediction-based screening, that's 5 hits per month: 60 confirmed candidates over the year. At 50% hit rate with prediction-based pre-screening (a realistic improvement in our experience for property-sensitive composition spaces with decent training data), you get 10 hits per month from the same budget: 120 confirmed candidates. Or equivalently, you reach the same 60 confirmed candidates in 6 months, freeing budget and calendar for the next phase of the program.
The doubling of hit rate isn't guaranteed. It requires that the property you're predicting has a genuine composition dependence learnable from available data, and that you have enough training data to build a model that's better than the researcher's unaided intuition. For mature chemistry spaces with existing literature data and internal measurement history, these conditions are often met. For genuinely novel chemistry without analogues in any database, prediction adds less value because there's nothing to train on.
Starting before you have "enough" data
The objection we hear most often: we don't have enough data to train a model. Our historical measurements are too few, too inconsistent, or not digitized.
The answer is that you rarely need as much data as assumed, especially if you're willing to use composition-to-property transfer from public DFT databases and literature-curated datasets as a starting point. A model that's 60% accurate at ranking candidates within your target chemistry space provides real value even if it's not publication-quality. You're not trying to publish a benchmark model. You're trying to filter a synthesis queue.
The relevant question is not "is my model good enough to be perfect?" but "is my model better than random?" If the answer is yes, the expected hit rate improves, and each confirmed measurement you add improves the model further. The workflow is self-reinforcing once you start, which is why waiting for a perfect dataset before beginning is the wrong frame. Start with what you have, accept that early predictions have wide uncertainty bounds, and accumulate measurements that incrementally sharpen the model's predictions in the regions that matter most to your program.
That's how the 70% failure rate moves over time. Not with a single methodological breakthrough, but with consistent application of the feedback loop: predict, prioritize, synthesize, measure, update. The information problem is tractable; it just requires committing to that loop from the beginning of a campaign rather than after several unsuccessful discovery cycles.