, ,

From CAD to AI: How Synthetic Data Is Changing Machine Vision Training

The algorithms behind AI inspection are advancing quickly. For many manufacturers, however, the harder problem is obtaining enough representative images to train them. Synthetic data is beginning to shift that work upstream, from the production line to the digital model.

The most persistent obstacle in industrial AI is not always the AI. Manufacturers can choose from an expanding range of deep-learning models, vision platforms and hardware. What they often cannot obtain quickly is the right training data. An inspection model needs to see more than a collection of flawless products photographed under ideal conditions. It must cope with the variation of a working line: changes in finish and colour, small shifts in position, reflections, shadows, new product variants and, most importantly, the defects that occur least often.

Those rare cases are usually the ones that matter most. A critical crack, missing component or assembly error may appear only occasionally, leaving teams to wait for examples, recreate faults by hand or begin with an incomplete dataset. New products pose an even more basic problem: the inspection system may need to be developed before production samples exist.

Manual labelling adds another cost. Specialists may have to classify images, draw boxes around components or trace defects at pixel level. Ambiguous quality criteria can also introduce inconsistency between annotators. Together, these factors have made data collection one of the longest stages in many AI vision projects. Synthetic data offers a way to reduce that dependency by generating training images in simulation, often using engineering assets a manufacturer already owns.

CAD moves beyond the design department

CAD models have traditionally described what a product should be: its geometry, dimensions and relationship to other components. They can now also become the foundation of a virtual inspection scene.

The model is placed in a simulated environment with digital cameras, lighting and backgrounds. Physics-based rendering uses ray tracing and material models to reproduce light across different surfaces. Polished metal, moulded polymer and transparent packaging each create different imaging conditions. This virtual scene can form part of a wider digital twin of the product, inspection cell or production line. Rather than building one visually perfect image, engineers can generate thousands of variations. Light intensity and direction can change; cameras can move; materials, backgrounds and object orientations can be adjusted. Image noise, occlusion and position can also be varied.

Engineer Working on Desktop Computer, Screen Showing CAD Software with Engine 3D Model, Her Male Project Manager Explains Job Specifics. Industrial Design Engineering Facility Office

This process is often described as domain randomisation. Its purpose is to prevent a model from becoming too dependent on one narrowly defined simulation. By training across a wide range of conditions, the system has a better chance of recognising the underlying product when it encounters real factory images.

The approach is already reflected in industrial software. NVIDIA Omniverse Replicator and Isaac Sim support synthetic-data generation and domain randomisation for perception and robotics, while Siemens has demonstrated the automatic creation of annotated images from 3D CAD data. These platforms are part of a broader move to reuse design and simulation data across the manufacturing lifecycle.

Building datasets before products exist

The immediate attraction is timing. If a usable CAD model exists, work on the vision application can begin before the first physical product or inspection station is available. Teams can render “golden” images representing acceptable products, then introduce controlled changes. Synthetic defects might include scratches, dents, contamination, damaged coatings, missing components or incorrect assembly. Their size, position and severity can be varied systematically, creating far more examples than a manufacturer could easily collect from early production.

The environment can change at the same time. A part can be viewed from different camera positions, under alternative lighting arrangements and against several backgrounds. It can be rotated, partly hidden or presented with neighbouring objects. This is particularly useful in robotics, where the same component may arrive in many poses rather than at a fixed point on a conveyor.

Because the simulation created the scene, it already knows what is in every image. It can produce bounding boxes, segmentation masks, depth information, defect classes and object poses automatically. That ground truth removes much of the manual annotation associated with real photographs and can be generated consistently at scale.

From Design Data to Deployed Inspection
CAD model → synthetic scene generation → lighting and material simulation → defect generation → automatic labelling → model training → validation on real images → deployment

Camera and lighting concepts can also be compared before hardware is installed. A proposed viewpoint may hide a critical surface, or reflections may overwhelm the feature being inspected. Finding those problems in simulation can influence the inspection cell while changes remain relatively inexpensive.

Training starts earlier

Synthetic datasets can feed many of the same applications as conventional image collections. Object-detection models can learn to locate and identify components. Instance segmentation can separate individual parts or map the precise area of a defect. Classification models can distinguish known fault categories, while anomaly-detection systems learn the appearance of acceptable products and flag departures from that baseline.

For robotic handling, synthetic data can also support pose estimation: determining where an object is and how it is oriented. That information enables robots to pick mixed or randomly positioned parts, align tools and plan inspection views. MVTec, for example, describes CAD-generated training data for deep 3D matching in bin-picking and robot-guidance applications.

Simulation does not make the system complete before production. It allows a manufacturer to reach commissioning with a working model and tested pipeline. Early real samples can then refine the system and focus data collection on identified weaknesses.

This could be particularly valuable for high-mix production, where frequent product changes make large, manually collected datasets impractical. It may also shorten deployment for greenfield projects by allowing vision development to run alongside mechanical and line engineering rather than after them.

The limits of simulation

Synthetic data works best when the relevant variation can be described. Geometry, position, lighting, component presence and camera viewpoint are relatively controllable. A missing fastener or rotated part can be generated accurately if its expected appearance is understood.

Other conditions are harder. Organic textures, fibres, liquids, transparent materials, complex reflections and irregular process damage may be difficult to reproduce convincingly. A rendered scratch may look plausible to a person while lacking the subtle visual characteristics of a real manufacturing fault. Simulation can also miss conditions no one anticipated, from vibration and dust to lens contamination or gradual tool wear.

This difference between simulated training images and factory images is the sim-to-real gap. It includes an appearance gap, where virtual materials or optics do not match reality, and a content gap, where the simulated dataset simply leaves out something that occurs in production. More photorealistic rendering can help with the first problem, but it cannot automatically identify every missing scenario. Real validation images therefore remain essential. Before deployment, a model must be tested against data captured with the actual camera, lens, lighting and production process. That validation set should include both normal variation and the difficult cases that determine whether the inspection is reliable.

For many projects, a hybrid dataset is likely to be more practical than an all-synthetic one. Simulation supplies volume, controlled variation and rare examples; a smaller set of real images anchors the model to the line. As production generates new conditions, those examples can be reviewed and added through continual learning, while the simulator is updated to cover newly discovered gaps.

That blended approach is explored in MVPro’s podcast episode, Synthetic Data in Practice with Dr Wilhelm Klein, Zetamotion. Klein describes “grounded synthetic data,” which begins with representative real images and expands the available variation synthetically. New production data and outliers continue to refine the system rather than being treated as evidence that training is finished.

The distinction matters. Synthetic data is sometimes presented as a replacement for factory data, but its more immediate role may be to make AI inspection projects viable sooner. It can reduce the need to wait for rare defects, move model development ahead of line commissioning and focus manual labelling on the images that add the most value.

A new role for engineering data

Synthetic data is changing where machine vision development begins. Instead of waiting beside a completed line for enough products and defects to appear, teams can start with CAD, simulation and a provisional understanding of the inspection task.

The result still depends on the assumptions built into the virtual scene, and factory validation remains the final test. But data generation is becoming part of inspection engineering rather than a separate exercise after production. As rendering, digital-twin and AI tools converge, CAD models are gaining a second operational life. They no longer describe only what should be manufactured. Increasingly, they can help train the systems that decide whether manufacturing has succeeded.

For more on the practical use of synthetic data when real defect examples are limited, listen to MVPro’s full episode: Synthetic Data in Practice with Dr Wilhelm Klein, Zetamotion.

Most Read

Related Articles

Sign up to the MVPro Newsletter

Subscribe to the MVPro Newsletter for the latest industry news and insight.

Name
Consent

Trending Articles

Latest Issue of MVPro Magazine

MVPro Media
Privacy Overview

This website uses cookies so that we can provide you with the best user experience possible. Cookie information is stored in your browser and performs functions such as recognising you when you return to our website and helping our team to understand which sections of the website you find most interesting and useful.