In a wind tunnel in Hampton, Virginia, a model wing glowed pink-purple under ultraviolet lights. High-speed cameras watched its surface change brightness as air pressed against it. The coating was a measuring instrument: pressure-sensitive paint, turning forces that cannot be seen into changes that a camera could record.
The experiment, described by NASA on October 2, used a standard research wing at Langley Research Center. It did not represent a particular aircraft. Its purpose was to help improve computer models by giving researchers something physical to compare their calculations against.
That relationship is becoming more complicated as artificial intelligence moves into engineering and science. A computer model can now supply the examples that teach another computer model. The student learns an approximation of the teacher, often much faster to run. But if the teacher leaves something important out, an apparently successful lesson can carry the omission forward.
On October 8, Siemens described software that learns from engineering simulations. A separate paper published that day followed AI systems from one simulated world into another, where their performance faltered. Together, they raise a question beneath the debate about how much training data AI needs: how much of the world does the data represent?
NASA did not announce AI training. Its experiment supplied measurements against which a computer model could be tested.
A Wing Under Pressure

The Benchmark Supercritical Wing before and during testing. NASA published the images on October 2, 2026.
NASA / Sarah Peak (left); NASA (right).
“The training data is work you already paid for,” Tyler Haggin, the chief operating officer of TrueInsight, wrote in the Siemens post.
He was describing Simcenter PhysicsAI, which can learn from completed structural and fluid simulations. Those files connect a proposed shape and conditions such as material properties with calculated results, including stresses spread across a three-dimensional object. The software can then predict results for another design. In the workflow Haggin describes, full simulation remains the reference for checking promising candidates.
The distinction matters because this is a different kind of education from reading engineering papers. A paper may explain what happened and why its authors think it happened. A collection of simulations can instead give a model thousands of worked examples, each with an input and a calculated answer. Their authority rests partly on decisions made before the AI arrived: which physical relationships were included, which could be simplified and which conditions the calculation was intended to cover.
More examples can fill in that imagined world without expanding its boundaries.
Consider a hypothetical factory model that represents the energy used while a furnace is running but omits the energy required to heat it up again. An AI trained to minimize consumption could learn to switch it off repeatedly. Within the model, the behavior would be economical. On a factory floor, it could waste energy. Further practice under the same rules would offer no reason to abandon the strategy. The system would have learned the assignment it was given.
That is an unusually awkward failure to detect from the training results alone. An error against the supplied answer is visible. An error shared by the answer and the learner can disappear from the score. The more faithfully the learner reproduces the flawed calculation, the better the training may look.
An October 8 paper in npj Artificial Intelligence described what happened when Serhii Aif and his colleagues trained ten AI agents to choose treatment actions in a simplified tumor simulation. Testing them in a more detailed simulation exposed failures.
Human analysis traced the discrepancy to growing cells pushing resistant cells toward the tumor’s edge, where growth conditions improved. Adding those effects to the training environment made performance more consistent. The team called the method Reinforcement Failing.
These were computational results, not improved patient outcomes. The more detailed simulation still omitted important biology. The authors said independent models and experiments were needed before clinical application.
There is a serious argument for leaving a great deal out. A simulation that reproduced everything would surrender much of the advantage of simulation. Scientists simplify because they want to isolate a relationship, repeat a calculation or investigate a possibility that would be difficult to arrange physically. An approximation can be exactly the right instrument for a narrow question.
AI adds another layer of approximation, but another opportunity as well. If a fast model allows many more candidates to be considered, the eventual choice may improve even when the finalists still undergo expensive checks. The economic comparison is between whole processes. A cheaper calculation that requires extensive repair may save little; a modest shortcut repeated often can matter considerably. Neither result follows from the prediction speed alone.
An earlier example, reported by UCLA on September 14, illustrates that bargain. A team working with Caltech and NVIDIA trained a model on simulations of hydronium, a positively charged molecule. It learned to predict responses to laser pulses, helping researchers find a sequence that would produce a desired state.
Against a reinforcement-learning approach with the same available controls, the method cut the reported time for generating a sequence from roughly ten hours to ten to twenty minutes. The comparison concerned a specific computational setup.
“For problems like this, we want models that learn the structure of the underlying physics itself,” Prineha Narang, the UCLA professor who led the team, said in the university’s account.
The qualifier matters. A model built for one molecular system can be useful without solving an unrelated problem. Replacing a slow calculation is already a consequential job; it does not require claiming to replace the experiment.
Physical evidence presents its own difficulties. A real experiment can still provide the wrong comparison if its conditions differ from the situation a model is supposed to describe.
In a September 28 report from MIT, researchers confronted that problem with a wind turbine only 15 centimeters across. At Princeton, they put it in a pressurized wind tunnel. Denser air helped the small machine reproduce important aspects of the flow around much larger turbines. Over several weeks, the team tested changes in alignment and operation, helping validate a predictive model. This was engineering model validation, not an AI training study.
Marcus Hultmark, the Princeton professor involved in the work, described the obstacle in the MIT account: “if you have nothing to compare them against, it’s very difficult to advance the field.” The comparisons themselves had required building the right experimental conditions.
That creates a limit on what abundant simulated data can buy. Computing can make more examples available, but multiplying examples from the same assumptions does not independently test those assumptions. A comparatively small physical experiment may settle a question that a much larger training run simply repeats. Its importance comes from the uncertainty it removes, which is not necessarily proportional to the size of the resulting file.
This also complicates the idea that science faces a single shortage of training data. One project may need broader coverage of designs it already knows how to simulate. Another may have ample simulated examples and little evidence about an unfamiliar operating condition. The first can benefit from more computation. The second may have to wait for equipment, expertise and a measurement that takes considerably longer to obtain than a prediction.
The computer could make its prediction in advance. At Langley, the cameras recorded what the air actually did to the wing, including anything the prediction might have missed.

