Automatic Recognition of Sand Foundry Defects Using Ultrasonic Testing and Convolutional Neural Networks

In the modern manufacturing industry, metal components produced by casting methods are widely used in automobiles, aerospace structures, and heavy machinery. During the casting process, internal defects such as porosity, shrinkage cavities, and cracks may form inside the component. These internal flaws directly compromise the structural integrity and service life of the final product. Therefore, reliable inspection and classification of sand foundry defects is a critical quality control step. Ultrasonic testing (UT) is one of the most commonly applied nondestructive testing (NDT) techniques for detecting internal defects. However, the interpretation of ultrasonic signals and images relies heavily on the experience of human operators, which limits the efficiency and consistency of the inspection. To overcome this limitation, my research aims to develop an intelligent algorithm based on deep learning that can automatically identify sand foundry defects from ultrasonic A-scan images. In this article, I present a systematic study covering ultrasonic detection principles, image preprocessing, neural network based recognition, and a novel adaptive convolutional kernel improvement for convolutional neural networks (CNN). Experimental results demonstrate that the proposed CNN architecture achieves a recognition accuracy above 90%, significantly outperforming traditional BP neural networks. The adaptive kernel strategy further improves the learning performance and test accuracy, confirming the effectiveness of the proposed approach for ultrasonic defect recognition in castings.

Introduction and Research Background

Castings are produced by pouring molten metal into a mold and allowing it to solidify. This process is preferred for its low cost, ability to form complex geometries, and suitability for large components. Nevertheless, defects are inevitably introduced during solidification, such as gas pores, shrinkage voids, inclusions, and hot tears. Among these, internal defects are particularly dangerous because they are invisible from the surface. In the field of NDT, ultrasonic testing has become the most popular method for internal defect inspection because it is safe, portable, and highly sensitive to volumetric and planar discontinuities. Although ultrasonic testing can effectively reveal the presence and position of internal defects, the final decision about defect type and severity is usually made by skilled inspectors. This manual process is subjective, time-consuming, and prone to fatigue-related errors.

With the rapid development of artificial intelligence, especially deep learning, there is a great opportunity to automate the recognition of sand foundry defects from ultrasonic data. Convolutional neural networks, in particular, have demonstrated remarkable success in image classification tasks. They can automatically learn hierarchical features from raw images without the need for hand-crafted descriptors. In this context, my thesis focuses on the entire pipeline of ultrasonic defect detection: acquiring ultrasonic images from test blocks, preprocessing the images to suppress background clutter, training both BP neural networks and CNNs on the processed images, and comparing their recognition performance. During the experiments, I observed that the size of the convolution kernel significantly influences the CNN performance. This observation led me to propose an adaptive convolution kernel method based on a defect factor, which is validated through a series of simulations.

The remainder of this article is organized as follows. First, I describe the fundamental principles of ultrasonic testing and the experimental setup used to acquire defect samples. Next, I detail the digital image preprocessing procedures, including binarization, cropping, edge detection, and morphology-based signal extraction. Then, I present the design and training of BP neural networks for defect recognition. After that, I introduce the CNN architecture, explain the dropout regularization technique, and report the recognition results. Finally, I propose and evaluate the adaptive convolution kernel optimization method. The article concludes with a summary and an outlook for future research.

Ultrasonic Testing of Castings

Ultrasonic Wave Properties

Ultrasonic waves are mechanical waves with frequencies above the audible range (20 kHz). In solid media, they propagate through the vibration of particles. Depending on the direction of particle displacement relative to the propagation direction, the waves can be classified as longitudinal waves, transverse waves, surface waves, or plate waves. Longitudinal waves, where particle displacement is parallel to the propagation direction, are the most commonly used in ultrasonic testing because they can travel through solids, liquids, and gases. The velocity c, frequency f, and wavelength λ are related by:

$$ \lambda = \frac{c}{f} = cT $$

where T is the period. The acoustic pressure p is proportional to the amplitude of the wave and determines the echo height displayed on the ultrasonic instrument. The acoustic intensity I represents the energy flux per unit area per unit time. These parameters are essential for quantitative assessment of defect size and location.

Basic Principle of Pulse-Echo Ultrasonic Testing

The pulse-echo method is the most widely used ultrasonic testing technique. A transducer emits a short ultrasonic pulse into the test object. When the pulse encounters a discontinuity, such as an internal defect or the back wall, part of the energy is reflected back to the transducer. By measuring the time delay between the transmitted pulse and the received echo, the depth of the reflector can be calculated. The amplitude of the echo provides information about the size and orientation of the reflector. The basic principle is illustrated in Figure 1, where T denotes the initial pulse, F is the defect echo, and B is the back-wall echo.

In my experiments, a digital ultrasonic flaw detector (model HS-600) was used together with straight and angle beam probes. The test blocks were steel castings with deliberately introduced artificial defects. Figure 2 shows photographs of the test blocks used in this study. A total of 110 sound (no-defect) images and 680 defective images were collected. Since the detector did not support real-time image transmission, all images were captured manually and then exported through the data management software. This manual data acquisition limited the total number of samples but provided a realistic dataset for training and testing the neural networks.

There are several common display modes for ultrasonic signals: A-scan, B-scan, and C-scan. The A-scan mode shows the echo amplitude as a function of time, which is the most direct representation. In the A-scan image, the horizontal axis is the propagation time and the vertical axis is the amplitude. The presence of a defect echo between the initial pulse and the back-wall echo indicates an internal discontinuity. My research adopts A-scan images as the input data for defect recognition.

Image Preprocessing and Edge Detection

Digital Image Representation

The raw images acquired from the ultrasonic flaw detector are color images with resolution 482×642 pixels and stored in RGB format. Since the actual content is grayscale (black and white), storing them as RGB is inefficient. Therefore, the first preprocessing step is to convert the RGB images to grayscale using the following weighted sum:

$$ Y = 0.2989R + 0.5870G + 0.1140B $$

This conversion reduces the image matrix from a three-dimensional array to a two-dimensional one. The grayscale image contains 256 gray levels, which can be further reduced to a binary image by thresholding. Binarization not only diminishes the data size but also simplifies the subsequent processing steps. In MATLAB, I used the function graythresh to obtain the optimal threshold based on the maximum inter-class variance method (Otsu’s method). The algorithm maximizes the between-class variance of the foreground and background pixels. The formulas used in Otsu’s method are:

$$ \omega_0 = \frac{N_0}{M \times N}, \quad \omega_1 = \frac{N_1}{M \times N} $$
$$ \mu = \omega_0 \mu_0 + \omega_1 \mu_1 $$
$$ g = \omega_0 (\mu_0 – \mu)^2 + \omega_1 (\mu_1 – \mu)^2 = \omega_0 \omega_1 (\mu_0 – \mu_1)^2 $$

where N0 and N1 are the numbers of pixels below and above the threshold, respectively, and M×N is the total number of pixels. By maximizing g, the threshold is automatically determined.

After binarization, the image still contains irrelevant components such as grid lines from the device interface, alphanumeric text, and warning lines. These elements would interfere with the defect recognition algorithm. To remove them, I first cropped the image to the region containing the ultrasonic waveform. The cropped region is shown in Figure 3. Then, I applied edge detection operators to analyze the contours of the waveform. In my experiments, I compared three classical edge detection operators: Canny, Sobel, and Laplacian of Gaussian (LoG). The Canny operator uses a multi-stage algorithm with Gaussian smoothing, gradient computation, non-maximum suppression, and thresholding. The Sobel operator approximates the gradient magnitude by convolving the image with two kernels:

$$ G_x = \begin{bmatrix} -1 & 0 & 1 \\ -2 & 0 & 2 \\ -1 & 0 & 1 \end{bmatrix}, \quad G_y = \begin{bmatrix} -1 & -2 & -1 \\ 0 & 0 & 0 \\ 1 & 2 & 1 \end{bmatrix} $$

The gradient magnitude is computed as:

$$ G = \sqrt{G_x^2 + G_y^2} $$

The LoG operator uses the second derivative and is more sensitive to noise, so Gaussian smoothing is applied first. The LoG kernel can be represented as:

$$ \nabla^2 f = f(x-1,y) + f(x+1,y) + f(x,y-1) + f(x,y+1) – 4f(x,y) $$

The results of edge detection are shown in Figures 4–6. Although edge detection gives a good outline of the waveform, it does not completely remove the background grid. The grid consists of isolated dots, while the ultrasonic waveform is a connected curve. This distinction led me to use morphological operations and connected component analysis.

Morphological Processing and Signal Extraction

Mathematical morphology provides a powerful tool for binary image analysis. The basic operations are erosion, dilation, opening, and closing. Dilation expands the boundaries of foreground objects, while erosion shrinks them. Opening is erosion followed by dilation, and closing is dilation followed by erosion. In binary images, these operations can be defined using set theory. Let A be the image and B a structuring element. The dilation of A by B is:

$$ A \oplus B = \{ x \mid (\hat{B})_x \cap A \neq \emptyset \} $$

Erosion is defined as:

$$ A \ominus B = \{ x \mid (B)_x \subseteq A \} $$

Opening and closing are given by:

$$ A \circ B = (A \ominus B) \oplus B $$
$$ A \bullet B = (A \oplus B) \ominus B $$

In my experiment, I applied these operations to the binarized and cropped images. The results indicated that erosion and opening can remove the dotted grid but may damage the thin waveform line. To preserve the continuous waveform, I used connected component labeling. The algorithm labels each connected component and computes its area. Since the ultrasonic waveform forms the largest connected component in the region of interest, I can keep only the components with area greater than a certain threshold. In MATLAB, this is implemented as follows:

L = bwlabeln(BW, conn);
S = regionprops(L, 'Area');
bw2 = ismember(L, find([S.Area] >= P));

where P is the area threshold. The final extracted waveform is shown in Figure 7. The background is completely removed, leaving only the ultrasonic A-scan signal. This extracted image is then used as the input for the neural network classifiers.

BP Neural Network for Sand Foundry Defect Recognition

Fundamentals of Artificial Neural Networks

Artificial neural networks are computational models inspired by biological neurons. A single neuron performs a weighted sum of its inputs, adds a bias, and applies an activation function. The mathematical model is:

$$ u_k = \sum_{i=1}^{n} \omega_{ki} x_i $$
$$ y_k = f(u_k + b_k) $$

where xi are the inputs, ωki are the weights, bk is the bias, f is the activation function, and yk is the output. Common activation functions include the step function, piecewise linear function, Sigmoid function:

$$ f(x) = \frac{1}{1 + e^{-x}} $$

and the hyperbolic tangent:

$$ f(x) = \tanh(x) = \frac{e^{x} – e^{-x}}{e^{x} + e^{-x}} $$

Recently, the Rectified Linear Unit (ReLU) has become popular:

$$ f(x) = \max(0, x) $$

The learning process of a neural network adjusts the weights based on a set of training samples. In supervised learning, the network output is compared with the desired output, and the error is used to update the weights. The Backpropagation (BP) algorithm is a gradient descent method that propagates the error backward from the output layer to the hidden layers. For a multi-layer perceptron, the weight update rule is:

$$ \omega_{ij}(t+1) = \omega_{ij}(t) + \eta \delta_j x_i $$

where η is the learning rate and δj is the error term for neuron j.

BP Network Design and Experiments

I implemented several BP networks using MATLAB to classify ultrasonic images into two categories: with sand foundry defects and without defects. The input layer size was determined by the image dimensions. For the original cropped images, the input layer had 300×500 = 150,000 nodes, which was computationally heavy. Later, I compressed the images to 30×30 pixels, reducing the input dimension to 900. The hidden layer contained 50 neurons, and the output layer had either 2 nodes or 1 node. For the two-node output I used [1 0] for defect-free and [0 1] for defective. For the single-node output, I used 0 for defect-free and 1 for defective. The activation function was originally the Sigmoid function, but it caused numerical issues during training. I replaced it with a modified ReLU function:

$$ f(x) = \max(0, x) \cdot \frac{x}{|x|} $$

This function preserves the sign and magnitude of positive values while maintaining nonlinearity. The training process used 90% of the dataset for training and 10% for testing. The error threshold was set to 0.01 and the maximum number of iterations to 2000. Table 1 summarizes the performance of the BP networks.

Network Configuration Training Convergence Test Accuracy
Output layer: 2 nodes, original image Converged slowly, error stuck around 0.1 66.67%
Output layer: 1 node, original image Faster convergence, error still decreasing 77.8%
Output layer: 1 node, compressed image (30×30) Better final error 88.89%

The results show that the single-output BP network is more effective than the two-output version. Image compression not only reduced computation time but also improved test accuracy. However, the recognition performance remained insufficient for industrial application. The limitation of BP networks lies in their fully connected structure, which does not exploit the spatial structure of images. This motivated me to adopt convolutional neural networks.

Convolutional Neural Network for Ultrasonic Defect Detection

Architecture of CNN

A convolutional neural network is a specialized type of neural network designed for grid-like data, such as images. It consists of three main ideas: local receptive fields, shared weights, and spatial downsampling (pooling). These concepts allow the network to learn translation-invariant features.

The convolution operation for a 2D image f and a kernel g is defined as:

$$ z(x,y) = \sum_{m=0}^{M-1} \sum_{n=0}^{N-1} f(x+m, y+n) \cdot g(m,n) $$

where M and N are the kernel dimensions. During convolution, the kernel slides across the image, and each element-wise multiplication produces a feature map. Multiple kernels are used to extract different features.

After the convolutional layer, a pooling layer reduces the spatial size of the feature maps. Common pooling methods include max pooling and average pooling. Max pooling selects the maximum value in each local region, while average pooling computes the mean. For a 2×2 pooling region with stride 2, the output size is reduced by half in both dimensions. This downsampling reduces computational load and provides a degree of translation invariance.

The final part of a CNN is typically one or more fully connected layers. The output layer often uses the Softmax activation function for multi-class classification. The Softmax function converts a vector of raw scores into probabilities:

$$ S_j = \frac{e^{x_j}}{\sum_{i=1}^{K} e^{x_i}} $$

where K is the number of classes.

Dropout for Overfitting Prevention

Since the ultrasonic dataset is relatively small (790 images in total), overfitting is a serious risk. Dropout is a regularization technique that randomly deactivates a fraction of neurons during training. For each training iteration, a hidden unit is kept with probability p or set to zero with probability 1-p. The output of a dropout layer is:

$$ y = r \cdot (W^T x) $$

where r is a binary vector sampled from a Bernoulli distribution with parameter p. During testing, no dropout is applied, but the weights are scaled by p to maintain consistency. Dropout prevents complex co-adaptations between neurons and improves generalization.

CNN Design and Results

I constructed a shallow CNN with one convolution layer, one pooling layer, one fully connected layer, and a Softmax output layer. The convolution layer used 20 convolution kernels of size 5×5. The pooling layer used 2×2 average pooling with stride 2. The fully connected layer had 100 neurons with the hyperbolic tangent activation function. The output layer had 2 nodes for defect-free and defective classes. The input images were the preprocessed binary images. The training set contained 90% of the samples and the test set contained 10%. The error threshold was set to 0.01, and the network was trained for up to 5000 iterations. The error curve during training is shown in Figure 8. The error decreased steadily and exhibited a clear downward trend, indicating that the CNN effectively learned the distinguishing features of sand foundry defects.

The final test accuracy reached 90%. In some repeated experiments, the CNN achieved 100% accuracy after sufficient training. Table 2 compares the performance of BP and CNN on the same dataset.

Method Input Size Test Accuracy
BP (single output, original) 150,000 77.8%
BP (single output, compressed) 900 88.89%
CNN (5×5 kernels) 150,000 90%
CNN (3×3 kernels, adaptive) 150,000 100%

These results demonstrate that CNN is significantly superior to BP networks for ultrasonic defect recognition. The CNN automatically learns relevant features directly from pixels, whereas the BP network treats each pixel independently and cannot exploit spatial correlations.

Adaptive Convolution Kernel Exploration

During the experiments, I noticed that changing the size of the convolution kernel affected the CNN performance. To investigate this effect, I conducted a series of tests with kernel sizes ranging from 3×3 to 13×13. Table 3 lists the average training time and test accuracy for each kernel size.

Kernel Size Training Time (s) Test Accuracy (%)
3×3 398.11 94
5×5 349.54 88
7×7 318.52 76
9×9 294.18 65
11×11 273.75 61
13×13 240.69 62

As the kernel size increases, the training time decreases because the convolution operation becomes faster with fewer multiplications. However, the test accuracy declines significantly. This is intuitive because a large kernel may over-smooth the fine details of the ultrasonic waveform, losing the subtle features that differentiate defects from background. Smaller kernels, such as 3×3, can capture more localized information and achieve better recognition accuracy.

Motivated by this observation, I proposed an adaptive convolution kernel approach based on a defect factor. The defect factor quantifies the proportion of foreground (waveform) pixels in the training images. Let the image be f(x,y) of size m×n. For a dataset of A images, the defect factor Kd is defined as:

$$ K_d = 1 – \frac{1}{A} \sum_{i=1}^{A} \frac{\sum_{x=1}^{m} \sum_{y=1}^{n} f_i(x,y)}{m \times n} $$

Since the waveform occupies a small portion of the image, Kd is close to 1. To select the kernel size, I multiply Kd by the image dimensions and round to the nearest integer. In my dataset, this calculation yielded a kernel size of 3×3. The adaptive kernel was then used to train the CNN again. The training error curve shown in Figure 9 exhibits even better convergence than the 5×5 kernel. The test accuracy improved to 100% on the test set, as shown in Figure 10. These results confirm that the adaptive kernel based on the defect factor is beneficial for the CNN performance.

It is worth noting that smaller kernels increase the computational cost because they require more convolution operations. In practical applications, one must balance accuracy and speed. Nevertheless, the adaptive kernel method provides a data-driven way to choose the kernel size instead of relying on manual tuning.

Conclusion and Future Work

In this research, I developed a complete framework for automatic recognition of sand foundry defects using ultrasonic A-scan images. The main contributions can be summarized as follows:

First, an effective image preprocessing pipeline was established. By combining binarization, cropping, edge detection, and connected-component morphology, the ultrasonic waveform was successfully extracted from the raw images, eliminating background clutter that would otherwise degrade classification performance.

Second, a BP neural network was implemented and tested for defect recognition. The experiments showed that a single-output BP network with compressed images achieved 88.89% accuracy. Although this was better than a two-output network, it was still insufficient for reliable industrial use.

Third, a CNN with one convolutional layer, one pooling layer, and one fully connected layer was designed. The CNN achieved 90% accuracy on the test set and showed superior learning capability compared to BP. This demonstrates the strong potential of deep learning for ultrasonic defect detection.

Finally, I discovered that the convolution kernel size significantly affects CNN performance. To address this, I proposed an adaptive kernel method based on a defect factor, which automatically selects the kernel size from the training data. Experimental results confirmed that the adaptive kernel improved both training convergence and test accuracy.

There are several directions for future work. The current CNN only distinguishes defective from non-defective samples. It could be extended to classify different types of sand foundry defects, such as porosity, shrinkage, and cracks. Additionally, the network may be further optimized to estimate defect size and location. In real-world applications, the detection environment may change, and labeled data may be scarce. Therefore, semi-supervised or unsupervised learning methods, such as autoencoders, could be explored to enhance the adaptability of the system. Finally, more advanced CNN architectures and transfer learning techniques could be investigated to achieve even higher accuracy and robustness.

In conclusion, this study proves that convolutional neural networks, especially with adaptive kernel design, are highly effective for recognizing internal defects in castings from ultrasonic images. The proposed method offers a promising solution for automating the inspection of sand foundry defects, improving the efficiency and reliability of quality control in the casting industry.

Scroll to Top