[ad_1]
Replica description
The canvas was realized using linen applied on a wooden frame (15 × 20 cm), and the creation of the replica followed a dedicated procedure performed by a professional artist (now deceased), who had extensive experience in painting.
The support material for the mock-up is linen (cf. Fig. 2), which is a high-quality base material commonly used for historic paintings on canvas. The canvas was then evenly coated with a water-based polymer glue to provide better adhesion and stability. On top of this, a ground layer (cf. Fig. 2), made of white pigment and adhesive, was applied, which not only provides a smooth surface for painting but also supports the drawing process that follows. On this ground layer, the initial pencil drawing was completed, which includes a reproduction of the pentimento found in the original drawing and the artist’s signature on the mock-up. Next, the paint layer is described. In order to protect the painting layer (cf. Fig. 2), the surface was evenly coated with a layer of varnish (cf. Fig. 2). This protective coating prevents the pigments from being damaged by the elements and prolongs the life of the work. Finally, the entire work was mounted in a wooden frame measuring 15 × 20 cm, ensuring its structural stability and providing a traditional form of display.

From top to bottom, these are the varnish, paint layers, ground, size, and canvas. The charcoal drawing is indicated on the right.
A portion of Mylar inserts and biological attack are shown in Fig. 3a. Mylar is a polyester film made from stretched polyethylene terephthalate. This material was used to simulate defects in the canvas layers. Areas of bio-attack defects caused by microorganisms are located on the back of the canvas, which changes over time in the real world.

Fabrication of the replica: a the back of the replica with bio-attack defects, b the initial pencil drawing on the ground layer, c the linen canvas, d the preparation of the Mylar insert, e using Mylar to make defect A, f allowing the linen canvas layer to dry, g the ground layer with defect A and defect B, h the final replica, i the back of the final replica.
More specifically, after the initial pencil drawing on the ground layer, A (4.5 × 1.0 cm) and B (4.0 × 1.0 cm) defects, using Mylar (flexible transparent polyester film), were placed between the size layer (i.e., the glue layer) and the ground layer as shown in Fig. 3b–g. Defects A and B are located between the layers, with A being thicker and B being thinner. Another defect (3.5 × 4.1 cm), i.e., a sponge, was glued to the lower left corner of the back of the painting, as shown in Fig. 3i. The sponge was wetted with 80 mL of water using a syringe to simulate the rising damp effect. Figure 3h shows the image after the painting was finished; defect D, a Mylar insert in the linen canvas is not visible here and was therefore not analyzed thereafter.
Replica geometrics
Establishing a good geometry model is essential for reliable numerical simulations, as it determines heat-transfer paths and the distribution of temperature across the model. To closely replicate the actual structure, the geometric model was first developed in Autodesk AutoCAD with precise X, Y, and Z configurations, simulating the canvas and frame of the original artifact. This model (cf. Fig. 4) was then imported into COMSOL Multiphysics software to perform detailed heat-transfer simulations.

a Front view of the replicated geometric model. b Rear view of the replicated geometric model. c Side view of the replicated geometric model. The defect regions highlighted by red boxes indicate the simulation regions of interest (ROIs). The x, y, and z axes represent the coordinate system used in the model.
Once imported, meshing was applied within COMSOL to optimize mesh density and element types, thus reducing numerical instability and ensuring accurate thermal analysis. Environmental and boundary conditions were configured to mirror the environmental conditions, enabling a temperature distribution that reflects real-world scenarios. Since data scarcity often limits defect detection in heritage conservation, COMSOL Multiphysics provided a comprehensive simulation framework by incorporating various fields—such as heat transfer, electromagnetism, or structural mechanics—via the finite element method (FEM). This approach generated high-quality, synthetic datasets that supported robust training for deep learning models.
The numerical simulation results confirmed experimental data trends, providing a theoretical foundation for subsequent image processing and defect analysis. By combining the simulated temperature distribution data with deep-learning feature extraction, high-quality inputs were obtained, addressing data scarcity and enhancing defect detection reliability.
Replica simulation
Because SLT uses natural solar radiation as the excitation source, a heat-exchange model is established to define the boundary conditions for COMSOL simulations. It is important to distinguish between the forward thermal model and the infrared camera radiometric chain. In this work, COMSOL is used as a heat-transfer forward model that outputs the time-dependent surface temperature field \(T(x,y,t)\) of the multilayer canvas replica.
The following describes the three steps for establishing the SLT model. To establish the SLT model, solar heating at the illuminated surface is modeled first, followed by radiative and convective heat exchange with the environment, and finally, a time-dependent representation of irradiance and the numerical setup are specified.
In this study, the sample receives heat through solar irradiation at the illuminated surface to reproduce the heating process under insolation conditions. Let \(S\left(t\right)\) denote the incident solar irradiance \((W/{m}^{2})\). The absorbed solar heat flux applied to the front surface is written as:
$${q}_{{\rm{abs}}}\left(t\right)={{\rm{\alpha }}}_{{sol}}S(t)$$
(1)
where \({qabs}(t)\) is the absorbed short-wave heat flux \((W/{m}^{2})\), and \({{\rm{\alpha }}}_{{sol}}\) is the effective solar absorptivity (or absorption efficiency) of the surface; Note that the illuminated area is only required when computing the total absorbed power \({Q}_{{abs}}(t)={q}_{{abs}}(t)\) \({A}_{{illum}}\), and should not be included in the heat-flux definition.
After heating, the sample dissipates heat to the environment mainly through long-wave radiation and convection. The net long-wave radiative heat flux at the sample surface is modeled by the Stefan–Boltzmann law:
$${q}_{{\rm{rad}}}=\varepsilon \times \sigma \times ({T}_{{\rm{amb}}}^{4}-{T}^{4})$$
(2)
where \(\sigma =5.670* {10}^{-8}\,{\rm{W}}/{{\rm{m}}}^{2}{{\rm{K}}}^{4}\) is the Stefan–Boltzmann constant; \({T}_{{amb}}\) is the ambient temperature in \({K}\); \(T\) is the surface temperature of the sample in \(K\); and \(\varepsilon\) is the surface long-wave emissivity of the surface. We emphasize that \(\varepsilon\) governs long-wave thermal radiation exchange, whereas solar absorption in Eq. (1) is governed by \({\alpha }_{{sol}}\).
Convective heat exchange is modeled as:
$${q}_{{conv}}=h({T}_{{amb}}-T)$$
(3)
where \(h\) is the convective heat-transfer coefficient \((W\cdot {m}^{-2}\cdot {K}^{-1})\).
To parameterize the day-scale evolution of irradiance and to generate different solar-loading levels for simulation-based augmentation, a compact diurnal model may be used:
$$S(t)={S}_{\max }\times \sin (\frac{t-{t}_{{sr}}}{{t}_{{day}}})\,t\epsilon [{t}_{{sr}}\,,{t}_{{ss}}]$$
(4)
where \({S}_{\max }\) is the maximum solar radiation intensity at the noon hour of the day; \({t}_{{sr}}\) and \({t}_{{ss}}\) are sunrise and sunset times, and \({t}_{{day}}={t}_{{ss}}-{t}_{{sr}}\).
This expression is an irradiance (boundary-input) model and does not represent the infrared camera response. For short acquisition windows (e.g., 10 s in this study), \(S(t)\) can be approximated as constant and set to the measured irradiance associated with the corresponding experimental session.
Based on Eqs. (1)–(4), the absorbed solar heat flux is applied as a boundary heat input on the illuminated front surface, while convection and long-wave radiation are applied on the exposed surfaces. By defining the thermophysical parameters of each layer (e.g., thermal conductivity and specific heat capacity, cf. Table 1) a complete transient heat-transfer model is constructed in COMSOL to compute \(T(x,y,t)\).
Replica experimental
This section describes the experimental setup, followed by the data acquisition protocol. In this study, the thermal imaging camera used was the FLIR T1020, which is equipped with UltraMax resolution mode to improve imaging accuracy with temperature measurements Fig. 5a. The camera’s emissivity (ε) was set to 0.95, the distance between the camera and the object was ~165 cm and the elevation angle of the lens was ~27°, as shown in Fig. 5b. The FLIR T1020 has a high temperature sensitivity and a wide range of temperatures, allowing it to accurately capture small temperature changes on the surface or subsurface of the test object. Particularly in cultural-heritage inspection, the device demonstrates excellent adaptability to detect subtle defects, such as cracks or splitting.

a Schematic diagram of the infrared thermal imaging experimental acquisition device. b Photograph of the actual experimental setup.
In order to analyze the effectiveness of the IRT technique under different environmental conditions and to ensure that potential defects can be effectively detected at different temperatures, humidity, and time intervals, four sets of IRT data were acquired in this study. These four sets of data were recorded at different temperatures, humidity and image acquisition rates, covering different time periods from morning to afternoon, taking into account the effects of environmental changes on the imaging results, see the following section.
A small sample has always been a significant challenge during the training process of deep-learning methods. To overcome this, the present study uses the COMSOL software and FEM method to build a simulation model of the painting on canvas with defects. The model was constructed using a combination of CAD and COMSOL interactive modeling functions, and the resulting simulated structure is shown in Fig. 4.
The mock-up has a total thickness of 9 mm, a surface coating thickness of 0.5 mm, a length of 200 mm and a width of 150 mm. To reproduce the properties of the real sample as closely as possible, the mock-up is made of multiple layers of solid material arranged according to the actual material distribution. Table 1 lists the material properties.
Two types of known defects were introduced on the front and back of the replica simulation model: splitting as Mylar inserts and bio-attack defects. To increase the diversity of the training data, this study designed thirty samples with different defect configurations, including different defect locations and sizes. Figure 6 shows four sets of defect configurations on the simulation model. The size and position of the defects in each simulation are different.

The yellow arrows in defect A and B indicate the placement and orientation of defects in the simulation model, with the defect size represented by solid square lines.
In the IRT simulation, solar radiation was used as the main heat source. A solar radiation heat flux was applied to the front surface of the sample through the solid heat conduction module, and a gradual change was simulated. A convective heat-transfer coefficient of 10 W/(m²\(\cdot\)K) was applied to the remaining surfaces. The simulation lasted 10 s, with a time step of 0.02 s, in order to accurately simulate the heat transfer and heat dissipation behavior under solar radiation.
In this way, 15,000 simulated infrared images were generated. However, in order to enhance the generalization and robustness of the model training, 1000 randomly selected images were generated from the 15,000 simulated infrared images, and the images were scaled and deformed to augment the number of images to 1500 for model training. Among these images, 90% were used for the training set, and the remaining 10% were used for the validation set.
The training phase was performed on the Windows 10 operating system using the deep learning framework PyTorch 1.12.0. For hardware configuration, an Intel Xeon E5-2670 processor, an NVIDIA GTX 3060 graphics card, and a 2TB hard drive were used. The defect detection and analysis experiment refers to multiple open-source codes and mechanisms, including the ECA mechanism30,31 and the YOLOv7 network structure32. The algorithm is implemented based on Python 3.7.13 and executed in the PyCharm integrated development environment. After multiple optimizations, the optimal hyperparameters of the YOLOv7 network were determined: the learning rate was 0.001, the optimizer was Adam, and the training was conducted for 1000 epochs
Methods
In this study, we employ the YOLOv7 model32 as the detection backbone, enhanced with attention mechanisms integrated into the feature extraction stage. For a detailed explanation of the YOLOv7 architecture and its loss functions, can refer readers to Wang et al. (2023).
The localization loss measures the position difference between the ground truth box and the predicted box through the mean square error (\({MSE}\)), and the formula is:
$${L}_{\mathrm{loc}}=\frac{1}{N}\mathop{\sum }\limits_{i=1}^{N}{\left({y}_{i}-{\hat{y}}_{i}\right)}^{2}$$
(5)
where \(N\) is the number of prediction boxes, \({y}_{i}\) is the position of the true box, and \({\hat{y}}_{i}\) is the position of the prediction box.
The classification loss usually uses the Cross-Entropy Loss, which is expressed as follows:
$${L}_{\mathrm{cls}}=-\frac{1}{N}\mathop{\sum }\limits_{i=1}^{N}\mathop{\sum }\limits_{c=1}^{C}{y}_{\mathrm{ic}}\log \left({\hat{y}}_{\mathrm{ic}}\right)$$
(6)
where \(N\) is the number of samples, C is the number of categories, \({y}_{{\rm{ic}}}\) is the true category label (0 or 1), and \({\hat{y}}_{{\rm{ic}}}\) is the predicted category probability.
Finally, the multi-task loss function of YOLOv7 can be expressed as:
$$L={\lambda }_{{\rm{loc}}}{L}_{{\rm{loc}}}+{\lambda }_{{\rm{cls}}}{L}_{{\rm{cls}}}+{L}_{{\rm{conf}}}$$
(7)
where \({L}_{{\rm{conf}}}\) is the confidence loss and \({\lambda }_{{\rm{loc}}}\) and \({\lambda }_{{\rm{cls}}}\) are weight coefficients that balance the influence of different loss terms. With this multi-task loss function design, YOLOv7 can simultaneously focus on accurate target positioning and classification during training, thereby improving overall detection performance.
Additional modules support network optimization: the efficient layer aggregation network module (cf. Fig. 7), managing gradient paths to facilitate effective learning; the Transition_Block (cf. Fig. 7), which reduces computational complexity through down-sampling; and RepConv (cf. Fig. 7), enhancing gradient diversity to boost generalization across feature maps.

The network comprises a backbone, ELAN-based blocks, an SPPCSPC module, an FPN-like feature-fusion pathway, transition blocks, and three detection heads. Arrows indicate feature flow, and dashed boxes indicate the main functional modules.
In recent years, attention mechanisms, particularly spatial attention and channel attention, have become increasingly important in the development of deep-learning models for computer vision. The evolution of attention mechanisms can be broadly divided into several stages. Early studies focused on spatial attention, which aims to identify the most informative spatial regions within an image. Subsequently, channel attention mechanisms were introduced to adaptively emphasize important feature channels and enhance the representation capability of CNNs. Representative developments include the squeeze-and-excitation network (SENet), which performs channel-wise feature recalibration, and the convolutional block attention module (CBAM), which combines spatial and channel attention to improve performance across various vision tasks.
In this study, channel attention mechanisms, including SE, coordinate attention (CA), and ECA, were incorporated into the feature-extraction stage of YOLOv7 to enhance feature selection and suppress environmental noise in infrared thermal images, which is essential for accurate defect detection in cultural-heritage artifacts. By adaptively emphasizing informative channels while suppressing less relevant responses, these mechanisms enhance the discriminative capacity of convolutional features.
Among these approaches, SENet introduces a channel attention mechanism that adaptively estimates the importance of feature channels, thereby strengthening feature representation for tasks such as classification and detection. SENet consists of three main operations: squeeze, excitation, and recalibration.
The input feature map \(X\epsilon {{\mathscr{R}}}^{(H\times W\times C)}\) is subjected to global average pooling (GAP), which compresses the spatial information of each channel to generate a global description. The global information of each channel is obtained by aggregating spatial information across the feature map:
$${z}_{c}=\frac{1}{H\times W}\mathop{\sum }\limits_{i=1}^{H}\mathop{\sum }\limits_{j=1}^{W}{X}_{{ijc}}\,,\,c=1,2,\ldots ,C$$
(8)
where \({z}_{c}\) denotes the global information feature of the \(c\mathrm{th}\) channel; H and W are the height and width of the feature map; and \(C\) is the number of channels.
In the squeeze step, the spatial response of each channel is compressed into a scalar descriptor by GAP, allowing the network to capture global contextual information.
In the excitation stage, channel weights are generated using a two-layer fully connected network. First, the global feature vector \(z=[{z}_{1},{z}_{2},\ldots {z}_{C}]\) is subjected to a dimensionality-reduction operation and is then restored to the original number of channels to generate the weights of each channel. The specific formula is:
$${s}_{c}=\sigma \left({W}_{2}\delta \left({W}_{1}\times z\right)\right)$$
(9)
where \({W}_{1}\,\epsilon \,{{\mathscr{R}}}^{\frac{C}{r}\times C}\) is the reduced weight matrix that reduces the feature space to \(\frac{C}{r}\); \({W}_{2}\,\epsilon \,{{\mathscr{R}}}^{(C\times \frac{C}{r})}\) is the weight matrix used to recover the reduced features to the original channel dimension. The ReLU activation function, denoted as \(\delta\), is defined as \(\delta \left(x\right)=\max \left(0,x\right)\); \(\sigma \left(x\right)=\frac{1}{1+{e}^{-x}}\) ensures the weights are in the range [0,1].
The purpose of the excitation process is to capture the nonlinear dependencies between channels and generate channel weights through the fully connected layer \(s\), which are then used to adjust the relative importance of each channel.
The generated channel weights, \(s=\left[{s}_{1},\,{s}_{2},\ldots ,{s}_{C}\right]\) are applied to the original feature map \(X\) on each channel to accomplish channel recalibration:
$${\widetilde{X}}_{c}={s}_{c}\times {X}_{c}$$
(10)
where \({s}_{c}\) is the weight of the cth channel, and \({X}_{c}\) is the original feature map of the cth channels. Through this weighting process, the network emphasizes informative channels and suppresses irrelevant responses, thereby improving the robustness of feature representation.
CA extends channel attention by encoding long-range channel dependencies together with positional information along the spatial dimensions.
Similar to SENet, the CA mechanism begins with GAP to compress the input feature map \(X\epsilon {{\mathscr{R}}}^{(H\times W\times C)}\). However, instead of only compressing the spatial dimensions globally, CA divides the spatial compression along the height and width dimensions separately to preserve positional information:
$${z}_{c}^{h(i)}=\frac{1}{W}\mathop{\sum }\limits_{j=1}^{W}{X}_{i,j,c}\,,\,{z}_{c}^{w(j)}=\frac{1}{H}\mathop{\sum }\limits_{i=1}^{H}{X}_{i,j,c}$$
(11)
where \({z}_{c}^{h}\) is a feature vector obtained by global pooling of the input feature map \({X}\) in the height dimension, representing the compressed information of each channel in the height dimension; \({z}_{c}^{w}\) denotes the aggregation of information in the width dimension and represents the global information extracted in the width direction; \({X}_{i,j,c}\) represents the individual element (i.e. pixel value or feature value) in feature map \(X\) located in the ith row, jth column, and cth channel.
This step helps the model capture both channel and spatial information simultaneously.
In the next step, CA generates weights for each channel, but this process is now guided by the positional information captured along the height and width dimensions. The weight generation function is similar to the excitation step in SENet, using fully connected layers to generate the attention weights for each channel. However, CA introduces separate transformations for the height and width:
$${s}_{c}^{h}=\sigma \left({W}_{2}^{h}\delta \left({W}_{1}^{h}\times {z}_{c}^{h}\right)\right)$$
(12)
$${s}_{c}^{w}=\sigma \left({W}_{2}^{w}\delta \left({W}_{1}^{w}\times {z}_{c}^{w}\right)\right)$$
(13)
where \({s}_{c}^{h}\) indicates the channel weight generated in the height dimension; \({s}_{c}^{w}\) represents the channel weight generated in the width dimension; \({W}_{1}^{w}\) and \({W}_{2}^{w}\) are the weight matrices in the width dimension; \(\delta\) is the ReLU activation function; \({W}_{1}^{h}\) and \({W}_{2}^{h}\) represent the weight matrices in the height dimension.
Finally, the recalibrated feature maps \({\widetilde{X}}_{c}\) are obtained by applying the channel weights separately for height and width:
$${\widetilde{X}}_{i,j,c}={s}_{c}^{h}{X}_{i,j,c}\times {s}_{c}^{w}{X}_{i,j,c}$$
(14)
This step ensures that both channel-wise and spatial attention are considered, enabling the model to focus on the most informative regions and channels of the input feature map.
Although the fully connected layers used in SENet improve channel selectivity, they also increase computational cost. ECA addresses this limitation by avoiding dimensionality reduction and replacing the fully connected transformation with a lightweight 1D convolution.
Indeed, a key innovation of ECA is the introduction of the 1D convolution with an adaptive kernel size, which is determined by the number of feature channels \(C\). This design enables the model to capture local channel dependencies without relying on fully connected layers. The kernel size \(K\) is adjusted according to the number of channels, and its value is determined by the following formula:
$$K=\varphi \left(C\right)={\left|\frac{{\log }_{2}(C)}{\gamma }+\frac{b}{\gamma }\right|}_{\mathrm{odd}}$$
(15)
where \(\psi \left(C\right)\) is a function that adaptively selects the kernel size based on the number of channels \(C\). \({\left|\cdot \right|}_{\mathrm{odd}}\) denotes that the closest odd integer to the calculated value is taken. The parameter \(\gamma\) controls how sensitive the kernel size is to changes in the number of channels \(C\), whereas the translation factor \(b\) fine-tunes the final kernel size after applying the logarithmic function. This adaptive strategy enables the model to capture local cross-channel interactions efficiently.
Unlike the two fully connected layers of SENet, ECA directly recalibrates through 1D convolution. The recalibrated feature map is as follows:
$${s}_{c}=\sigma \left({\rm{Conv}}1D\left({z}_{c}\right)\right)$$
(16)
where \({z}_{c}\) is the global information descriptor of the cth channel obtained by GAP and \(\sigma\) denotes the sigmoid activation function; the result is a set of weights \({s}_{c}\), which are applied to the original feature map \({X}_{c}\) to adjust the importance of each channel. Finally, the obtained\(\,{s}_{c}\) is used in Eq. (10) to obtain \({\widetilde{X}}_{c}\). By simplifying the channel recalibration process, ECA significantly reduces the computational complexity of the model while still effectively enhancing the feature representation.
The YOLOv7 model integrates three attention mechanisms—SE, CA, and ECA—in the feature extraction stage. These mechanisms, when integrated, allow YOLOv7 to selectively focus on multi-dimensional features, crucial for enhancing detection in complex object detection scenarios. Although the loss function remains consistent with the original YOLOv7, this study examines which attention mechanism (SE, CA, or ECA) optimally improves network performance for infrared thermal imaging. Identifying the most effective mechanism is essential for advancing automated detection accuracy in cultural-heritage conservation.
[ad_2]
Source link

