In my research, I focus on the automatic detection of internal defects in sand foundry defect specimens using ultrasonic testing combined with advanced neural network architectures. The motivation arises from the fact that cast metal components produced by sand casting methods often contain internal flaws such as porosity, shrinkage cavities, and inclusions. These internal discontinuities significantly affect the mechanical properties of the final products. Traditional ultrasonic testing relies heavily on manual interpretation of A-scan or B-scan images, which is subjective, time-consuming, and inefficient. To address this challenge, I developed a systematic approach that integrates digital image preprocessing, feature extraction, and deep learning–based classification. Specifically, I compare the performance of a backpropagation neural network (BPNN) with a convolutional neural network (CNN) for the identification of sand foundry defect signatures in ultrasonic images. Additionally, I propose an adaptive convolution kernel strategy to optimize CNN performance. Experimental results demonstrate that the proposed CNN achieves over 90% classification accuracy, significantly outperforming the traditional BP network. The adaptive kernel method further improves learning efficiency and accuracy, confirming its potential for real-world industrial applications.
1 Introduction
Sand casting is one of the oldest and most widely employed manufacturing processes for producing metallic components. It is favored for its low cost, ability to form complex geometries, and suitability for large-scale production. However, sand foundry defect formation is an inherent risk during the casting process. Defects such as gas porosity, shrinkage cavities, sand inclusions, and hot tears can severely degrade the quality and reliability of cast parts. Therefore, non-destructive testing (NDT) methods are essential for quality control. Among various NDT techniques, ultrasonic testing is particularly attractive because of its high penetration capability, sensitivity to internal flaws, safety, and portability. In the context of sand foundry defect detection, ultrasonic waves can interact with internal discontinuities and produce characteristic echoes that reveal the presence, location, and approximate size of defects.
In conventional ultrasonic inspection, the A-scan display presents a one-dimensional waveform representing echo amplitude versus time. B-scan and C-scan displays provide two-dimensional cross-sectional images. However, interpretation of these images is usually performed by human experts, whose judgment can be subjective and inconsistent. To improve efficiency and objectivity, I explore automated recognition of sand foundry defect patterns from ultrasonic images using machine learning and deep learning algorithms. The primary contributions of this work include:
- Development of a complete image preprocessing pipeline to extract ultrasonic signal waveforms from raw device screenshots.
- Design and implementation of a BP neural network for defect classification, including an investigation of network output configurations.
- Design and implementation of a convolutional neural network with dropout regularization for the same classification task.
- Comparison of BP and CNN performance, showing that CNN is significantly superior for sand foundry defect recognition.
- Proposal of an adaptive convolution kernel method based on a defect factor to optimize CNN performance.
The rest of this paper is organized as follows. Section 2 describes ultrasonic testing principles and data acquisition. Section 3 details the image preprocessing and edge detection methods used to isolate the ultrasonic signal. Section 4 introduces the neural network fundamentals and the BP neural network experiments. Section 5 presents the convolutional neural network architecture, training procedure, and results, along with the adaptive kernel optimization. Finally, Section 6 concludes the work and suggests future directions.

2 Ultrasonic Testing Principles and Data Acquisition
2.1 Fundamentals of Ultrasonic Waves
Ultrasonic waves are mechanical waves with frequencies above the human audible range, typically greater than 20 kHz. They can propagate through solid, liquid, and gaseous media, although transverse waves cannot propagate in fluids. The key parameters describing a wave are the wavelength $\lambda$, frequency $f$, and propagation velocity $c$, which are related by:
$$ \lambda = \frac{c}{f} = cT, $$
where $T$ is the period of the wave. In steel, the longitudinal wave velocity is approximately $5.96 \times 10^{3} \ \text{m/s}$, while the transverse wave velocity is approximately $3.23 \times 10^{3} \ \text{m/s}$. For sand foundry defect inspection, longitudinal waves are commonly used because of their ease of generation and high sensitivity to volumetric discontinuities.
The acoustic pressure $p$ is proportional to the amplitude of the received echo and is directly displayed as the vertical deflection on an ultrasonic instrument screen. The acoustic intensity $I$ is the energy flow per unit area per unit time. These quantities are essential for interpreting the strength of reflections from internal defects.
2.2 Pulse-Echo Method
The most widely used ultrasonic testing method for sand foundry defect detection is the pulse-echo method. A piezoelectric transducer emits a short ultrasound pulse into the test object. When the wave encounters an internal discontinuity, part of its energy is reflected back to the transducer. The time delay between the initial pulse and the received echo indicates the depth of the reflecting interface. In the absence of defects, only the initial pulse $T$ and the bottom echo $B$ appear on the display. When a defect exists, an additional echo $F$ appears between $T$ and $B$, as illustrated in the conceptual diagram below.
The depth $d$ of a defect can be calculated from the time delay $t_d$ using the relation:
$$ d = \frac{c \cdot t_d}{2}, $$
where $c$ is the wave velocity in the material. This principle is fundamental for locating sand foundry defect inclusions in cast components.
2.3 Equipment and Data Collection
For my experiments, I used a digital ultrasonic flaw detector (HS-600 series) equipped with straight and angled probes. The test specimens used were steel castings and alloy cast blocks with artificially manufactured internal defects. During data acquisition, I manually placed the probe on the specimen surface with an ultrasonic couplant (glycerin) to ensure efficient transmission of sound energy. The A-scan images displayed on the instrument screen were saved using the device’s data management software. Because the device did not support real-time image streaming, each image had to be captured and exported manually. In total, I collected 110 images of defect-free regions and 680 images containing internal defects. This dataset formed the basis for training and testing the neural network models.
The raw images contain not only the ultrasonic waveform but also background grid lines, digital readings, and system interface elements. These extraneous features must be removed before the images can be used for automatic recognition. The complete preprocessing procedure is described in the next section.
3 Image Preprocessing and Edge Detection
3.1 Digital Image Representation
A digital image is represented as a two-dimensional array $f(x,y)$, where $x$ and $y$ are spatial coordinates and the value of $f$ represents the intensity (or gray level) at each pixel. Color images are stored using three channels (red, green, blue), each requiring a separate matrix. To reduce computational complexity, I first converted the original RGB images to grayscale using the weighted formula:
$$ I_{gray} = 0.2989 R + 0.5870 G + 0.1140 B. $$
After conversion, each pixel has a single gray value in the range 0–255. Further reduction to binary images is achieved through thresholding, which greatly reduces the data volume and simplifies subsequent processing.
3.2 Binarization and Image Cropping
I applied Otsu’s method to automatically determine an optimal threshold $T$. The method maximizes the inter-class variance between foreground and background pixels. Given an image $I(x,y)$, the threshold $T$ is found by maximizing:
$$ g = \omega_0 \omega_1 (\mu_0 – \mu_1)^2, $$
where $\omega_0$ and $\omega_1$ are the proportions of foreground and background pixels, and $\mu_0$ and $\mu_1$ are the corresponding mean gray values. Binary images have only two gray levels, which makes morphological operations and connected-component analysis straightforward. After binarization, I manually cropped the region of interest containing the echo waveform to eliminate the device system interface and other irrelevant parts.
3.3 Edge Detection Operators
To isolate the ultrasonic waveform, I investigated three classical edge detection operators: Canny, Sobel, and Laplacian of Gaussian (LoG). The Sobel operator computes gradient approximations using two convolution kernels:
$$ G_x = \begin{bmatrix} -1 & 0 & 1 \\ -2 & 0 & 2 \\ -1 & 0 & 1 \end{bmatrix}, \quad G_y = \begin{bmatrix} -1 & -2 & -1 \\ 0 & 0 & 0 \\ 1 & 2 & 1 \end{bmatrix}. $$
The gradient magnitude is given by:
$$ G = \sqrt{G_x^2 + G_y^2}. $$
The LoG operator uses a second-order derivative and is defined as:
$$ \nabla^2 f(x,y) = f(x+1,y) + f(x-1,y) + f(x,y+1) + f(x,y-1) – 4f(x,y). $$
In my experiments, all three operators successfully extracted edges, but none of them completely removed the background grid and digital text. Therefore, I adopted a different strategy based on mathematical morphology and connected-component analysis.
3.4 Morphological Processing and Connected-Component Analysis
Morphological operations include dilation and erosion. Dilation expands bright regions and erosion shrinks them. Opening is erosion followed by dilation, while closing is dilation followed by erosion. These operations are defined as:
$$ A \oplus B = \{ x \mid (\hat{B})_x \cap A \neq \emptyset \}, $$
$$ A \ominus B = \{ x \mid (B)_x \subseteq A \}, $$
$$ A \circ B = (A \ominus B) \oplus B, $$
$$ A \bullet B = (A \oplus B) \ominus B. $$
In the binary ultrasonic images, the useful signal is a continuous curve, while the background consists of isolated dots and text characters. By labeling connected components and computing their areas, I could identify the largest connected region, which corresponds to the ultrasonic waveform. The algorithm proceeds as follows:
- Compute connected components using
bwlabeln. - Calculate the area of each component with
regionprops. - Keep only components with area above a threshold $P$ using
ismember.
This procedure yielded clean images containing only the ultrasonic signal, as confirmed visually. Table 3.1 summarizes the effects of different processing steps on the image quality and background removal.
| Method | Background removal | Signal preservation | Computation load |
|---|---|---|---|
| Binarization only | Poor | High | Low |
| Sobel edge detection | Moderate | Moderate | Medium |
| Canny edge detection | Moderate | High | Medium |
| Morphology + connected components | Excellent | High | Medium |
The final preprocessed images were resized to a smaller dimension (e.g., 30×30 pixels) to reduce the input size for the neural networks. This resizing preserved the essential features while significantly decreasing the number of input nodes and thus the training time.
4 Backpropagation Neural Network for Defect Recognition
4.1 Neural Network Fundamentals
An artificial neuron is the basic processing unit of a neural network. It computes the weighted sum of its inputs and applies an activation function. The output of a neuron is expressed as:
$$ u_k = \sum_{i=1}^{n} \omega_{ki} x_i, \quad y_k = f(u_k + b_k), $$
where $\omega_{ki}$ are the weights, $x_i$ are the inputs, $b_k$ is the bias, and $f(\cdot)$ is the activation function. Common activation functions include the sigmoid, hyperbolic tangent, and rectified linear unit (ReLU). In my BP experiments, I found that the standard sigmoid function sometimes caused numerical instability during training. Therefore, I replaced it with a modified ReLU activation defined as:
$$ f(x) = \max(0, x) / \max(1, |x|), $$
which preserves the non-linearity while avoiding extremely large outputs.
4.2 BP Algorithm
The backpropagation algorithm updates the network weights by minimizing the mean squared error between the network output and the desired output. For an output layer with $q$ neurons, the error function is:
$$ E = \frac{1}{2} \sum_{i=1}^{q} (d_i – yo_i)^2. $$
The weight update rules for the hidden-output and input-hidden connections are:
$$ \omega_{hy}(t+1) = \omega_{hy}(t) + \eta \delta_y(t) ho(t), $$
$$ \omega_{ih}(t+1) = \omega_{ih}(t) + \eta \delta_h(t) x(t), $$
where $\eta$ is the learning rate, $\delta_y$ and $\delta_h$ are the error gradients for the output and hidden layers, respectively. The delta terms are computed using the derivative of the activation function.
4.3 BP Network Design for Sand Foundry Defect Detection
I designed a BP network with one input layer, one hidden layer containing 50 neurons, and an output layer. Two configurations were tested: a two-node output layer using [1 0] for “no defect” and [0 1] for “defect”, and a single-node output layer using 0 for “no defect” and 1 for “defect”. The input layer size was initially equal to the number of pixels in the full-size preprocessed image (e.g., 300×500 = 150,000 nodes), but this led to excessive computational costs and poor convergence. To mitigate this, I resized the images to 30×30 pixels, reducing the input nodes to 900. Table 4.1 summarizes the BP network configurations and their results.
| Configuration | Input size | Hidden nodes | Output nodes | Training error (final) | Test accuracy (%) |
|---|---|---|---|---|---|
| Full image, 2-output | 150,000 | 50 | 2 | ~0.15 | 66.7 |
| Full image, 1-output | 150,000 | 50 | 1 | ~0.10 | 77.8 |
| Compressed image, 1-output | 900 | 50 | 1 | ~0.05 | 88.9 |
The training process was repeated multiple times with random initialization. The best representative training curves show that the single-output network with compressed images achieved faster convergence and lower final error. However, even with this improvement, BP networks suffered from instability and often converged to local minima. This limitation motivated me to adopt convolutional neural networks, which are better suited for image feature extraction and classification tasks.
5 Convolutional Neural Network for Ultrasonic Defect Detection
5.1 Convolutional Layer
A convolutional neural network employs local receptive fields, shared weights, and sub-sampling to extract spatial features from images. The convolution operation between an input image $f(x,y)$ and a convolution kernel $g(x,y)$ is defined as:
$$ z(x,y) = (f * g)(x,y) = \sum_{m=0}^{M-1} \sum_{n=0}^{N-1} f(m,n) g(x-m, y-n), $$
where $M$ and $N$ are the kernel dimensions. In practice, the kernel slides across the image with a certain stride, and the sum of element-wise products forms a feature map. The output size $O$ for an input of size $I$, kernel size $K$, padding $P$, and stride $S$ is:
$$ O = \left\lfloor \frac{I – K + 2P}{S} \right\rfloor + 1. $$
In my CNN architecture, I used a single convolutional layer with 20 convolution kernels of size 5×5. The activation function after convolution was the hyperbolic tangent. This layer extracted low-level features such as edges and local patterns that are discriminative for sand foundry defect identification.
5.2 Pooling Layer
Pooling (or sub-sampling) reduces the spatial dimensions of feature maps while preserving important information. I employed average pooling with a 2×2 filter and stride 2. The pooling operation for a region $\mathcal{R}$ is:
$$ y_{pool} = \frac{1}{|\mathcal{R}|} \sum_{(i,j) \in \mathcal{R}} x_{ij}, $$
for average pooling, or $y_{pool} = \max_{(i,j) \in \mathcal{R}} x_{ij}$ for max pooling. Pooling reduces the number of parameters and provides translation invariance. In my design, the output of the pooling layer was a set of 20 reduced feature maps.
5.3 Dropout and Overfitting Control
Since the dataset is relatively small, overfitting is a common risk. I applied dropout regularization to the fully connected layer. Dropout randomly sets a fraction $p$ of neurons’ outputs to zero during training, which prevents complex co-adaptations. The output of a dropout layer is:
$$ y = r \odot (W^T x), \quad r_i \sim \text{Bernoulli}(p), $$
where $\odot$ denotes element-wise multiplication. In my experiments, I set the dropout probability to 0.5 during training. This significantly improved the generalization performance of the CNN.
5.4 CNN Architecture and Training
The complete CNN architecture is listed in Table 5.1. The network consists of an input layer (30×30 grayscale image), a convolutional layer with 20 filters of size 5×5, a tanh activation, an average pooling layer (2×2), a fully connected layer with 100 neurons, a tanh activation, a dropout layer (p=0.5), and an output layer with 2 neurons using the softmax activation function for binary classification.
| Layer | Type | Dimension/Parameters | Activation |
|---|---|---|---|
| Input | Image | 30×30×1 | – |
| Conv1 | Convolution | 20 filters, 5×5, stride 1 | tanh |
| Pool1 | Average pooling | 2×2, stride 2 | – |
| FC1 | Fully connected | 100 neurons | tanh |
| Dropout | Dropout | p=0.5 | – |
| Output | Fully connected | 2 neurons | softmax |
The softmax function converts the output scores into probabilities:
$$ S_j = \frac{e^{z_j}}{\sum_{i=1}^{C} e^{z_i}}, $$
where $C$ is the number of classes. The network was trained using gradient descent with a learning rate of 0.01 and a maximum of 10,000 epochs. The training error decreased rapidly, as shown in the error curve during training. After training, I tested the network on the reserved test set, which comprised 10% of the total data. The test accuracy reached 90% on average, and in several repetitions reached 100% after sufficient convergence.
5.5 Comparison with BP Neural Network
Table 5.2 compares the performance of the optimized BP network and the CNN. The CNN demonstrates significantly better accuracy, stability, and generalization. The BP network often produced inconsistent results across different training runs, whereas the CNN consistently achieved high accuracy. This confirms that convolutional feature extraction is highly effective for recognizing sand foundry defect patterns in ultrasonic images.
| Model | Architecture | Test accuracy (%) | Convergence | Stability |
|---|---|---|---|---|
| BP (compressed) | 900–50–1 | 88.9 | Moderate | Low |
| CNN (5×5 kernel) | Conv-Pool-FC | 90.0 | Fast | High |
6 Adaptive Convolution Kernel Optimization
During my experiments, I observed that the size of the convolution kernel has a strong influence on CNN performance. To quantify this effect, I trained identical CNN architectures with kernel sizes ranging from 3×3 to 13×13. The average training time and test accuracy for each kernel size are summarized in Table 6.1.
| Kernel size | Average training time (s) | Average test accuracy (%) |
|---|---|---|
| 3×3 | 398.1 | 94 |
| 5×5 | 349.5 | 88 |
| 7×7 | 318.5 | 76 |
| 9×9 | 294.2 | 65 |
| 11×11 | 273.8 | 61 |
| 13×13 | 240.7 | 62 |
Smaller kernels generally yield higher accuracy, while larger kernels reduce computational cost but with a significant degradation in performance. This trend arises because smaller kernels are better at capturing fine-grained features that are crucial for distinguishing subtle differences between defect echoes and noise. Based on this observation, I proposed an adaptive convolution kernel method guided by a defect factor $K_d$, which measures the proportion of the ultrasonic signal area in the image. The defect factor is defined as:
$$ K_d = 1 – \frac{1}{A \cdot m \cdot n} \sum_{i=1}^{A} \sum_{x=1}^{m} \sum_{y=1}^{n} f_i(x,y), $$
where $f_i(x,y)$ is the binary value of the $i$-th image, $A$ is the number of training images, and $m \times n$ is the image size. The convolution kernel size is then chosen as:
$$ K_{size} = \lceil K_d \cdot \max(m,n) \rceil. $$
For the dataset used in this study, the computed kernel size was 3×3, which matches the best-performing fixed-size kernel. By using this adaptive scheme, the CNN automatically selects the optimal kernel size for the given data distribution. I further increased the number of training epochs to verify the effectiveness of the 3×3 kernel. The improved CNN achieved a training error curve that converged faster and reached a lower final error than the 5×5 kernel, and the test accuracy reached 100% in the final run. The results confirm that the adaptive kernel method offers a promising direction for optimizing CNN-based sand foundry defect detection systems, although the increased computational time due to smaller kernels must also be considered in real-world applications.
7 Conclusion and Future Work
In this research, I successfully developed an automated ultrasonic testing framework for the detection of sand foundry defect internal discontinuities in castings. The main conclusions can be summarized as follows:
- Image preprocessing using binarization, cropping, morphological operations, and connected-component analysis effectively isolates the ultrasonic echo signal from the complex background of the instrument screen.
- A BP neural network is capable of classifying ultrasonic images, but its performance is highly dependent on network configuration and suffers from instability. The best BP network achieved an accuracy of 88.9% after image compression.
- A convolutional neural network with a single convolutional layer, pooling, and dropout reached an average test accuracy of 90%, demonstrating that CNN is significantly more suitable for image-based defect recognition than traditional fully connected networks.
- The size of the convolution kernel greatly affects CNN performance. An adaptive kernel strategy based on the defect factor successfully selected the optimal kernel size (3×3) for the dataset, improving both convergence and test accuracy.
For future work, I plan to extend the approach to multi-class defect classification, such as distinguishing porosity, shrinkage cavities, and inclusions. Furthermore, I intend to investigate deeper CNN architectures and unsupervised feature learning methods to reduce the reliance on labeled datasets. Finally, the real-time deployment of the CNN model on embedded ultrasonic inspection devices would be a valuable industrial application.
In conclusion, the combination of ultrasonic imaging, digital image processing, and convolutional neural networks offers a robust and efficient solution for automatic sand foundry defect detection, which can significantly enhance quality control in metal casting industries.
