Casting Defect Detection via Deep Learning

Introduction and Motivation

As a researcher working on industrial quality inspection, I have focused my master’s thesis on the automated detection of casting appearance defects, particularly those found on automotive brake brackets. These components frequently exhibit surface imperfections such as dark holes, shallow pits, cracks, depressions, bulges, and notches. Such defects not only reduce the mechanical strength of the casting but also create loose connections between assembled parts, posing serious safety risks. Traditional inspection methods, based on manual visual inspection or classical non-destructive testing, are costly, inefficient, and heavily dependent on human expertise. In modern high-volume production lines, these limitations make them increasingly inadequate. Therefore, my research objective was to develop a robust and efficient detection algorithm based on deep learning that could accurately identify and localize sand foundry defects in real time.

The application of deep learning to sand foundry defects is not straightforward. Casting surfaces have a frosted texture with natural spots and grain, which can easily be confused with actual defects. Additionally, different defect categories share many visual similarities, while the same defect may vary greatly in size, shape, and intensity across samples. The challenges are further amplified by the need to process high-resolution images without losing critical details. In my work, I used a five-megapixel industrial camera to capture images of 2566×1940 pixels, preserving as much information as possible. Yet, feeding such large images directly into a convolutional neural network (CNN) is computationally prohibitive and often leads to poor performance because standard CNNs typically down-sample input images, which destroys the subtle features of sand foundry defects.

To overcome these obstacles, I designed a two-stage detection framework. The first stage employs a YOLO (You Only Look Once) network to detect and localize the regions of interest on the casting. The second stage uses an improved ResNet-50 network to classify each extracted region as either defect-free or containing a specific type of defect. This divide-and-conquer approach converts the difficult task of detecting tiny defects in a huge image into a more tractable problem: first find medium-sized casting regions, then analyze those regions for micro-defects. In the following sections, I describe the methodology, the improvements I made to the ResNet architecture, and the experimental results that demonstrate the effectiveness of the proposed system.

Task Analysis and Technical Foundation

Characteristics of Sand Foundry Defects

The casting defects I studied can be grouped into six categories plus the “no defect” class. Table 1 summarizes their typical appearances and root causes.

Defect Type Visual Appearance Typical Cause Impact
Dark Hole Deep, dark circular cavity Non-uniform solidification Reduces strength
Shallow Pit Light, shallow, often clustered spots Metal splashes Poor surface quality
Crack Linear fissure with irregular depth Thermal stress or mold collision Critical failure risk
Notch Missing material at an edge Mechanical impact Immediate rejection
Bulge Raised edge section Insufficient grinding Assembly problem
Depression Sunken edge section Over-grinding Loose connection

These defects are typically small, ranging from a few pixels to tens of pixels. For example, a dark hole may occupy only 10×10 to 50×50 pixels within a 2566×1940 image. Moreover, the surface texture of the casting itself contains spots and grain that mimic defects. This high inter-class similarity and intra-class variability make the identification of sand foundry defects a challenging pattern recognition problem.

Why Standard CNNs Fail

Classic CNN architectures such as AlexNet, VGG, and GoogLeNet are designed for large-scale image classification where the objects of interest occupy a significant portion of the image and are visually distinct. In those tasks, aggressive down-sampling to 224×224 or 256×256 works well because the key semantic features survive. For sand foundry defects, however, the defect is essentially a local texture anomaly. If I were to resize the full 2566×1940 image to 256×256, a 10-pixel defect would shrink to less than one pixel, becoming invisible. Even for larger defects, the down-sampling would blur critical boundaries and merge adjacent texture patterns, leading to misclassification.

Another issue is the imbalance between the defect size and the receptive field of the network. A single CNN that tries to detect tiny defects directly from a full-resolution image would need an extremely large feature map, which is computationally infeasible. Even modern object detectors like Faster R-CNN or SSD struggle with small targets when the input resolution is high. Therefore, a hierarchical approach is necessary.

Proposed Algorithm Overview

My proposed pipeline has two main stages, as illustrated in the flowchart below.

Stage 1: Region Detection with YOLO

The first stage locates all relevant regions on the casting. The brake bracket used in this study has six key areas: left platform, upper-left link, lower-left link, right platform, upper-right link, and lower-right link. These are the areas most prone to sand foundry defects. I trained a YOLO-v3 network to detect these six region categories. The input to YOLO is the full image resized to 416×416 pixels. Since the regions are relatively large and have distinctive geometric shapes, down-sampling to 416×416 does not lose their essential features.

The YOLO network divides the image into an S×S grid, and each grid cell predicts B bounding boxes, each with a confidence score and class probabilities. The output tensor is:

$$T = S \times S \times (B \times 5 + C)$$

where 5 denotes the four coordinates $(x, y, w, h)$ plus one confidence score, and $C$ is the number of class categories (here $C=6$). The confidence score for each box is defined as:

$$\text{Confidence} = \Pr(\text{Object}) \times \text{IOU}_{\text{pred}}^{\text{truth}}$$

where $\Pr(\text{Object})$ indicates whether the box contains an object, and $\text{IOU}$ measures the overlap between the predicted box and the ground truth. During inference, the network outputs a set of bounding boxes with class labels. I select boxes with confidence above a threshold, then apply non-maximum suppression to eliminate duplicates.

Once the regions are detected, I extract each region from the original high-resolution image. Since the region coordinates from YOLO are relative to the 416×416 resized image, I map them back to the 2566×1940 coordinate system. The extracted patches are then resized to 256×256 pixels. This patch size is sufficient to preserve most defect features because the defects reside entirely within one patch.

Stage 2: Defect Classification with Improved ResNet-50

The second stage classifies each 256×256 patch into seven classes: six defect types plus “no defect”. I chose ResNet-50 as the backbone because of its excellent balance between accuracy and computational cost. However, the default ResNet-50 uses ReLU activations and a single-stream architecture, which I found suboptimal for this task. I therefore introduced two key improvements.

First, I replaced the ReLU activation with a novel activation function called ASoftReLU. The standard ReLU is defined as:

$$f(x) = \max(0, x)$$

Although ReLU accelerates convergence and provides sparsity, it suffers from the “dying ReLU” problem: when a neuron’s input is consistently negative, its gradient becomes zero and it can never recover. To mitigate this, I combined ReLU with the Softplus function, $s(x) = \log(1+e^x)$, which has a smooth gradient everywhere. The ASoftReLU function is defined as:

$$f(x) = \max\left(x, a \log(1+e^x)\right), \quad 0 \le a \le 1$$

When $a=0$, ASoftReLU is identical to ReLU. When $a>0$, the negative part retains a small positive gradient, preventing dead neurons. The parameter $a$ controls the slope of the negative region. A larger $a$ gives a smoother transition and stronger gradient flow, but sacrifices sparsity. In my experiments, I tuned $a$ empirically and found that $a=0.3$ yields the best trade-off.

Second, I designed a multi-channel convolutional neural network architecture to enhance feature robustness. The idea is to augment the input image with multiple transformed versions and train several parallel ResNet branches on these versions. At the end, the predictions from all branches are averaged. The transformations include:

– Data augmentation: mirror flip, random rotation, random scaling, and random translation. These operations expand one input image into up to eight versions.
– Contrast enhancement: for each augmented image, I apply histogram equalization and imadjust (a grayscale intensity transformation). This produces three variants: the original, the histogram-equalized, and the imadjusted versions.

Thus, each channel of the multi-channel network receives a combination of these preprocessed images. The final prediction is the average of the softmax probabilities from all channels:

$$\hat{y} = \frac{1}{K} \sum_{k=1}^{K} \text{Softmax}(f_k(x_k))$$

where $K$ is the number of channels, $x_k$ is the $k$-th preprocessed image, and $f_k$ is the corresponding ResNet branch. This ensemble-like strategy reduces overfitting and improves the generalization ability of the network, especially when the training dataset is relatively small.

Experimental Setup

Dataset Preparation

For the region detection experiment, I collected 850 brake bracket samples. After discarding unusable images, I retained 625 original images. To expand the dataset, I applied random transformations (mirror, rotation, scaling, translation) to generate four images per original, yielding 2500 images in total. Among them, 2000 were used for training and 500 for testing. Each image was annotated with bounding boxes for the six casting regions.

For the defect classification experiments, I manually cropped 1600 region patches of size 256×256. The class distribution is shown in Table 2.

Class Dark Hole Shallow Pit Bulge Depression Crack Notch No Defect Total
Count 203 185 197 212 206 197 400 1600

I further augmented the training set to 6400 images using the same data augmentation operations, while the test set remained separate and unaugmented (1000 images). The test set contained 125 images of each defect class and 250 no-defect images.

Training Configuration

All experiments were conducted on a server with an Intel Core i7-4700 CPU, an NVIDIA RTX 2070 Ti GPU, and 32 GB RAM. I used TensorFlow with CUDA 9.1 and OpenCV 3.4 for image processing.

For the YOLO network, I adopted the standard YOLO-v3 architecture with Darknet-53 as the backbone. The input size was 416×416. I used a batch size of 10 and trained for up to 100 epochs. The initial learning rate was 0.001, with a decay schedule.

For the ResNet-50 classifier, I initialized the network with weights pre-trained on ImageNet and replaced the final fully connected layer with a new 7-class softmax layer. I used an exponential decay learning rate:

$$\eta = \text{lr} \cdot \text{decay\_rate}^{\text{global\_step} / \text{decay\_steps}}$$

with initial lr = 0.01, decay_rate = 0.1, and decay_steps = 20. The batch size was 50 images. The network was trained until the validation accuracy plateaued.

Evaluation Metrics

I evaluated the models using the following metrics:

– Precision: $P = \frac{TP}{TP+FP}$
– Recall: $R = \frac{TP}{TP+FN}$
– Mean Average Precision (mAP): average of per-class precision values.
– IOU (Intersection-over-Union): measures the overlap between predicted and ground-truth bounding boxes.
– FPS: frames per second for runtime evaluation.

For the classification tasks, I report two separate accuracies:
– Y/N accuracy: whether the patch is defective (binary).
– Category accuracy: exact defect type among the seven classes.

Experimental Results and Analysis

Experiment 1: YOLO Region Detection

Table 3 shows the detection accuracy on the test set at different training epochs.

Epoch Train Acc Test Acc IOU FPS
5 38.5% 33.4% 0.45 34
10 66.3% 61.3% 0.53 34
15 75.2% 67.2% 0.67 34
20 80.8% 74.2% 0.72 35
25 83.2% 75.5% 0.75 34
30 85.0% 80.8% 0.75 34
35 88.3% 84.2% 0.79 34
40 90.2% 86.6% 0.83 34
42 93.1% 87.0% 0.85 34
44 92.3% 87.8% 0.85 34
46 91.8% 88.5% 0.86 34
48 92.5% 89.0% 0.86 34
50 93.1% 88.0% 0.87 34
52 93.5% 88.4% 0.85 34
54 94.2% 87.7% 0.86 34

The best test accuracy of 89.0% was achieved at epoch 48. The detection recall was high: out of 3000 test region instances, 2933 were detected (97.8% detection rate). However, the correct classification of the region type was slightly lower. The per-category results are shown in Table 4.

Region Left Platform Upper-Left Link Lower-Left Link Right Platform Upper-Right Link Lower-Right Link Total
Ground Truth 500 500 500 500 500 500 3000
Detected 488 485 490 492 487 491 2933
Correct 444 426 427 453 429 432 2611
Precision 91.0% 87.8% 87.1% 92.1% 88.1% 88.0% 89.0%
Recall 88.8% 85.2% 85.4% 90.6% 85.8% 86.4% 87.0%

The lower accuracy on link regions compared to platforms is expected because links have more complex shapes and appear similar to each other. Nevertheless, the high detection rate ensures that almost every region is captured, which is sufficient for the subsequent defect classification stage.

Experiment 2: ASoftReLU Activation

I trained the ResNet-50 classifier with ASoftReLU activations for different values of the parameter $a$. The results are shown in Table 5.

Parameter a Epochs to Best Y/N Accuracy Category Accuracy
0 (ReLU) 63 91.3% 84.6%
0.1 63 91.6% 85.7%
0.3 65 92.1% 86.2%
0.5 65 91.8% 86.0%
0.7 66 91.7% 85.6%
0.9 67 91.5% 85.8%

The best results occur at $a=0.3$, improving the Y/N accuracy from 91.3% to 92.1% and the category accuracy from 84.6% to 86.2%. Moreover, the convergence speed with ASoftReLU is slightly better than with ReLU, as seen in the training curves. This confirms that the modified activation function helps avoid dead neurons and improves representational capacity for sand foundry defects.

Experiment 3: Multi-Channel Network

In this experiment, I evaluated multi-channel architectures with 1, 2, 4, 6, and 8 parallel branches. Each branch uses ASoftReLU with $a=0.3$. The training step count to convergence was approximately 340,000-350,000 across all configurations, indicating that the computational overhead of additional channels is not severe. Table 6 summarizes the accuracy results.

Channels Steps (×10³) Y/N Accuracy Category Accuracy
1 340 93.3% 86.4%
2 344 93.5% 87.1%
4 348 93.8% 87.9%
6 345 94.1% 87.5%
8 347 94.3% 88.2%

The 8-channel network achieves the highest performance. The detailed confusion information for the 8-channel model is shown in Table 7.

Predicted Dark Hole Shallow Pit Crack Notch Bulge Depression No Defect Total
Actual Dark Hole (125) 109 5 4 3 2 1 1 125
Actual Shallow Pit (125) 4 111 2 2 3 2 1 125
Actual Crack (125) 3 2 103 5 4 3 5 125
Actual Notch (125) 2 2 3 109 4 3 2 125
Actual Bulge (125) 2 3 2 4 108 3 3 125
Actual Depression (125) 1 2 2 5 3 101 11 125
Actual No Defect (250) 6 5 4 8 5 3 219 250

From Table 7, I observe that the “No Defect” class has the highest recall (96.4%), meaning that defect-free regions are rarely misclassified as defective. The “Depression” class is the most misclassified, often being confused with “No Defect” and “Notch”. This is plausible because depressions are subtle edge irregularities that may resemble normal grinding marks. Overall, the Y/N accuracy of the 8-channel system is 94.3% (computed as the sum of correct defect and no-defect predictions divided by 1000).

Comparison with Conventional Methods

To validate the superiority of my approach, I compared it with two traditional machine learning methods commonly used for sand foundry defects: KNN-based inspection and PCA+BP neural network. I also compared the performance of the original ResNet-50, the ASoftReLU version, and the final 8-channel version. The results are shown in Table 8.

Method Y/N Accuracy Category Accuracy
KNN 67.3% 55.2%
PCA+BP 75.2% 69.9%
ResNet-50 (baseline) 91.3% 84.6%
ResNet-50 + ASoftReLU 92.1% 86.2%
ResNet-50 + ASoftReLU + 8-channels 94.3% 88.2%

Deep learning methods substantially outperform the traditional approaches. My final system achieves a 94.3% Y/N accuracy and 88.2% category accuracy, which represents a 3.0% and 3.6% improvement over the baseline ResNet-50, respectively. These gains demonstrate the effectiveness of the proposed activation function and multi-channel architecture for detecting sand foundry defects.

Conclusion and Future Work

In this research, I developed a deep learning-based system for the automatic detection of casting appearance defects, specifically targeting sand foundry defects on automotive brake brackets. The system integrates a YOLO-v3 region detector with an improved ResNet-50 classifier. My main contributions are:

1. A two-stage pipeline that overcomes the challenges of high-resolution images and small defect sizes. By first locating casting regions and then analyzing each region independently, the system achieves high detection accuracy without compromising detail.

2. The introduction of the ASoftReLU activation function, which mitigates the dying neuron problem while preserving the advantages of ReLU. Experimental results show that $a=0.3$ provides the best trade-off, improving both convergence speed and classification accuracy.

3. A multi-channel convolutional neural network that combines data augmentation and contrast enhancement. This ensemble-like approach improves robustness against surface texture noise and inter-defect confusion. The 8-channel variant achieves the highest accuracy of 94.3% for defect presence and 88.2% for defect type.

Future work will focus on several aspects. First, the YOLO region detector still struggles with complex areas such as corners and connecting parts; I plan to refine the region detection accuracy by using a more powerful object detector or by incorporating semantic segmentation. Second, the classification accuracy for specific defect categories, especially depressions and notches, needs improvement. I intend to explore attention mechanisms and more advanced data augmentation strategies. Third, the current two-stage system runs at about 34 FPS for the YOLO part but is slower when both stages are combined. Optimizing the pipeline using model pruning and quantization could make it truly real-time in an industrial setting.

Overall, the proposed method demonstrates that deep learning can effectively automate the inspection of sand foundry defects, reducing cost and human error while maintaining high accuracy. I believe that this research provides a solid foundation for further developments in intelligent manufacturing quality control.

Scroll to Top