Results
Three refractory-alloy searches
Three runs were completed on A100 machines over the past week, all on the same design space of 64 near-equiatomic bcc alloys of Mo, Nb, Ta, W, V, Ti, Zr and Hf, each with 40 MACE evaluations and 10 Quantum ESPRESSO calculations, zero failed evaluations, and a passing acceptance report.
| run | controller graph | objective | screen | DFT | acceptance |
|---|---|---|---|---|---|
forager_rhea_v1 |
16,384 MaleCNS neurons | ΔH_mix, 0 K, unrelaxed | 40 | 10 | 205 checks, all pass |
forager_rhea_wholebrain_v1 |
163,972 neurons, 6,143,838 edges | ΔH_mix, 0 K, unrelaxed | 40 | 10 | 205 checks, all pass |
forager_rhea_wholebrain_T1500K_P5GPa_v1 |
163,972 neurons, 6,143,838 edges | ΔG_mix at 1500 K and 5 GPa | 40 | 10 | 207 checks, all pass |
The whole-brain graph is every typed neuron in MaleCNS v1.0 with at least five synaptic contacts per connection, reduced to its largest weakly connected component. Nothing was removed by a size cap. Extraction took 327 s and 5.6 GB of memory. A decision over 64 candidates took about 13 s and a learning update about 2.6 s on the VM's CPU.
This chapter reports what came out, in order of how much it should be trusted: the DFT numbers, then the screening model's accuracy, then what the controller did. The last part is the uncomfortable one.
What the alloys look like
At 0 K, unrelaxed. Ten alloys reached Quantum ESPRESSO (PBE, PAW pseudopotentials from pslibrary 1.0.0, 50/400 Ry cutoffs, 3×3×3 k-points, Marzari–Vanderbilt smearing 0.02 Ry). Four have negative same-structure mixing enthalpy: MoNbTaW −58, MoTaVW −58, MoTaV −44 and MoNbTaVW −36 meV/atom. Everything containing Zr or Hf is positive, between +62 and +91 meV/atom. TiVW sits at +19. That ordering is chemically unsurprising: group 5/6 refractory metals with close atomic radii mix nearly ideally in bcc; Zr and Hf are larger and not bcc at zero temperature, so forcing them into a rigid bcc supercell with a compromise lattice parameter costs energy.
At 1500 K and 5 GPa. Here the screening model first relaxed each cell and its atomic positions at 5 GPa, and DFT ran at that geometry. Relaxation kept every cell bcc: the median deviatoric cell strain was 0.8 % (largest 6.5 %) and no atom moved more than 0.31 Å. The QE ranking is now by ΔG_mix = ΔH_mix(P) − T·ΔS_conf:
| alloy | ΔG_mix | ΔH_mix(5 GPa) | T·ΔS_conf | DFT pressure at the relaxed cell |
|---|---|---|---|---|
| MoNbTiVW | −247 | −40 | 207 | 2.8 GPa |
| MoNbTaW | −233 | −54 | 179 | 2.4 |
| MoNbTaVW | −233 | −26 | 207 | 3.3 |
| MoTaVW | −223 | −44 | 179 | 3.6 |
| HfNbTiVW | −172 | +35 | 207 | 5.0 |
| HfNbTiW | −165 | +14 | 179 | 4.7 |
| MoTiWZr | −141 | +38 | 179 | 2.3 |
| TaVW | −133 | +8 | 142 | 4.4 |
| MoNbWZr | −129 | +51 | 179 | 1.5 |
| TaTiZr | −117 | +25 | 142 | 4.8 |
(meV/atom.) Two things to read off this table. First, the entropy term is three to five times larger than any enthalpy, so at 1500 K the ideal-solution model ranks by number of components first and by ΔH second; that is the model, not a discovery, and the book says so. Second, relaxation reduces the positive enthalpies substantially (MoNbWZr +91 → +51, HfNbTiW +63 → +14 meV/atom, comparing the two DFT runs on the same alloys) while barely changing the negative ones (MoNbTaW −58 → −54). Size-mismatched alloys gain a lot from local relaxation; well-matched ones do not. Both effects are physically expected.
The DFT pressure at the MACE-relaxed cells ranged from 1.5 to 5.0 GPa against a 5 GPa target. MACE's equation of state for the Mo/W-rich alloys is softer than PBE's, so its relaxed volumes are too large. Because the P·V term enters the alloy and its elemental references alike, the effect on ΔH_mix is second order, but the discrepancy is real and is recorded on every event.
None of these numbers is a property prediction. They are converged SCF energies of single 16-atom single-occupancy supercells. There is no special quasirandom structure, no vibrational entropy, no competing intermetallic phase and no oxidation, creep or ductility.
How good is MACE here
MACE-MP-0 (medium, float64, no dispersion) is the screening model. On the twenty alloys that both it and DFT evaluated at identical geometry:
| unrelaxed, 0 GPa (10 alloys) | relaxed at 5 GPa (10 alloys) | |
|---|---|---|
| Spearman rank correlation | 0.77 | 0.78 |
| mean absolute error | 70 meV/atom | 114 meV/atom |
| bias (MACE − DFT) | −60 meV/atom | −97 meV/atom |
| sign of ΔH agrees | 8 of 10 | 5 of 10 |
The rank order is decent and the absolute values are not. MACE is systematically too negative: it says MoNbTaW mixes at −153 meV/atom where DFT says −58, and after relaxation the gap widens to −180 against −54. On relaxed cells it calls five alloys stabilising that DFT calls destabilising. A least-squares line through the relaxed data gives DFT ≈ 0.28·MACE + 28 meV, which is to say the model's dynamic range is about three to four times too large for these systems.
That bias has a plausible cause. MACE-MP-0 was trained on Materials Project relaxations, which are dominated by ordered compounds; a random 16-atom bcc cell with four or five species and a compromise lattice parameter is far from that distribution, and the model over-rewards the local relaxations it finds there. The pure-element references relax to their equilibrium lattice parameters, where the model is accurate, so the error lands entirely on the alloy side of ΔH_mix.
For this project's purpose the conclusion is: MACE is usable to order candidates and is not usable as a reward. Its Spearman correlation of about 0.78 means a shortlist of the best ten by MACE will contain most of DFT's best ten. But rewarding a controller with MACE's ΔH (as the screen stage does) teaches it a landscape that is three times too steep and wrong in sign for half the borderline cases. Ideas for fixing that are at the end of the chapter.
What the controller did
The controller is a recurrent network on the measured graph. Each candidate alloy is encoded into a 20-dimensional feature vector, projected onto every neuron through a fixed random matrix, and the network settles for three steps of sparse recurrent dynamics. The candidate's score is the mean squared activity ("goodness") over all neurons after settling. Scores are turned into a softmax over unevaluated candidates at temperature 0.06, mixed with 15 % uniform exploration, and sampled. After the evaluator returns, a reward-minus-running-baseline signal scales a local Forward-Forward-style update: the best and worst replay records are replayed as a positive and a negative phase and each synapse moves by the local derivative of a per-neuron goodness loss. No gradient passes between neurons and no readout is trained.
The upper panel above is the finding. The probability with which the controller chose each alloy is identical across the three runs to within plotting precision, even though the first two runs used circuits ten times different in size and the third run received an almost entirely positive reward stream where the first two received a mostly negative one (lower panel). The forty alloys screened were the same set in all three runs, and the first twelve were the same alloys in the same order. The curve itself is 1/(number of alloys remaining): it is the shape of uniform sampling with a fixed random seed.
Re-scoring the whole candidate pool with the saved initial and final weights makes the same point without reference to the seed:
The score correlates with the norm of the candidate feature vector at r = 1.000; the recurrent circuit contributed nothing to it that the input norm did not already contain. Learning moved the weights a long way (‖Δw‖₂ = 111 on the ΔH run) and moved the choice distribution by at most 0.0015 in probability. The whole policy lives between 0.013 and 0.019 per alloy, where uniform is 0.0156.
In plain terms: in these three runs the fly brain did not choose the alloys. The seed did, with a slight bias toward Hf/Zr-rich compositions because they have the largest feature norms. MACE and Quantum ESPRESSO did all of the physics, and the DFT rankings above are exactly as valid as they would be for forty alloys picked at random, which is what they are.
This is not a subtle statistical effect and it does not need the ablation matrix to confirm it, though that matrix would have caught it earlier and should have been run before any A100 time was spent on whole-brain graphs. The cause is mechanical. Mean goodness over 164,000 neurons is a law-of-large-numbers quantity: it depends on the input drive's magnitude and almost nothing else, and a bounded local update on synapses whose incoming weights are renormalised to a fixed absolute sum can shift every neuron's goodness together (the ΔH run's scores all rose by 0.125) but cannot open a gap between candidates that the input did not already separate. Chapter 4 says the rule is not, by itself, a method for assigning credit; this is what that looks like in practice.
What should change before the next run
In order of importance:
- Give the controller a readout it can move. Score candidates from a designated output population rather than the whole-brain mean; for example the goodness of neurons that receive no direct input drive, so their activity is entirely a product of recurrent processing, or a fixed random subset of a few hundred neurons. Centre the scores across the candidate set before the softmax so only differences matter. Verify on the synthetic graph that a shuffled-reward control and a frozen-weights control produce different choice sequences from the learning arm; until they do, no biological claim is testable.
- Run the ablation matrix first, on the 16k graph, at MACE-only cost. Three seeds, rewired and generic-sparse controls,
randomandshuffled_rewardarms. This costs minutes, not A100-hours, and it is the experiment the whole project is built to do. - Do not reward with raw MACE ΔH. Either calibrate MACE against the DFT already in hand (a linear correction from the twenty shared points is a start), or reward on rank rather than value, or fine-tune MACE-MP-0 on the bcc supercells this project generates. The DFT budget spent so far is enough to start.
- Enlarge and diversify the design space when the controller works, not before. Sixty-four candidates with a 40-evaluation budget leaves almost nothing for a policy to exploit; sampling two thirds of the pool is an enumeration with extra steps. A few thousand compositions with non-equiatomic fractions, or a continuous composition parameterisation, would make selection matter.
- Treat the temperature term as what it is. Ideal configurational entropy will always favour the quinary. If temperature is to discriminate, it needs at least the quasi-harmonic vibrational term, and phase competition against Laves and σ phases, which is a different and larger calculation.
Replay the runs
| Run | Controller graph | Objective | Evaluations | Acceptance | |
|---|---|---|---|---|---|
forager_rhea_v1 | 16,384 MaleCNS neurons | ΔH_mix, 0 K, unrelaxed | 40 MACE · 10 QE | 205 checks, all pass | Open viewer → |
forager_rhea_wholebrain_v1 | 163,972 neurons · 6,143,838 edges | ΔH_mix, 0 K, unrelaxed | 40 MACE · 10 QE | 205 checks, all pass | Open viewer → |
forager_rhea_wholebrain_T1500K_P5GPa_v1 | 163,972 neurons · 6,143,838 edges | ΔG_mix at 1500 K, 5 GPa | 40 MACE · 10 QE | 207 checks, all pass | Open viewer → |