An animal model of radiological medical image reading: detection of lung abnormalities in multi-slice CT by pigeons (Columba livia)

59 min read Original article ↗

Introduction

Accurate detection of pathology is demanded of radiologists, but radiologists are human and therefore make errors. Some tasks, such as detecting pulmonary nodules in a lung CT, are a highly demanding visual search task with error rates that can rise to as high as 50% (Rubin et al., 2015; DiGirolamo et al., 2025). Artificial Intelligence has made significant progress in aiding radiologists’ evaluation of images; however, these computerized systems also make errors, sometimes missing an abnormality but more often mis-identifying healthy cases (Jin et al. 2023). Understanding how biological vision systems detect abnormalities remains important as the radiologists are the final determiners of abnormalities on these medical images. When examining a single medical image, the activity appears to follow typical visual search processing which has numerous feature-driven (bottom-up) and goal-driven (top-down) processes controlling observer efficacy (Wolfe 2021). These processes may reflect both explicit knowledge from instruction, like that received verbally in medical training, and also implicit visual processes accrued from experiences, like from images examined during clinical training (Smith et al. 2012).

In contrast to characterizations of this process as a type of visual search, implicit perceptual categorization may be a critical process in radiologists’ and pathologists’ visual analysis of medical images. A considerable amount of research across different labs and modalities has demonstrated that even brief presentations of a histology slide, mammograph, or MRI sections can be sufficient for human observers to correctly identify the category of an image (DiGirolamo et al. 2023; Evans et al. 2016; Treviño et al. 2020). In this experimental design, the image(s) are presented briefly (e.g., 48 ms/frame for MRI, 500 ms total for static mammograms or histology slides), sometimes with a visual “mask” to stop further visual processing, and the observers are required to report whether an abnormality was present. In some paradigms, the observer is also asked to report other information such as the location of a reported abnormality, their confidence in their rating, or the category of the abnormality (DiGirolamo et al. 2023). Critically, expert medical observers and perceptually trained novices are above chance categorizing these briefly presented images as abnormal.

Accurate responding after such a brief visual presentation, especially one followed by a visual mask, likely obviates explicit, language-based processing. Instead, implicit visual categorization guides the rapid processing and responding (Evans et al. 2013). The prevalence of implicit processing across imaging types as well as positive correlations between the observers’ expertise and implicit visual categorization skills (Evans et al. 2013; Treviño et al. 2020) suggests that this perceptual categorization process is an important component during effective medical image evaluation. Furthermore, Gandomkar et al. (2021) suggests that the implicit analysis of mammograms is independent of explicit judgment as performance on categorization was uncorrelated with their explicit responses. These data suggest additional perceptual information independent of explicit judgment that might improve the classification of cases (e.g., by machine learning analysis) if appropriately measured. Further, Evans et al. (2019) reports that the implicit signal can be used to predict a cancerous mass 3 years prior to when it can be explicitly visually detected. Thus, despite the emphasis placed on explicitly recognized features when training radiologists to detect abnormalities, improving the processing of implicit information may increase radiologists’ effectiveness in identifying patients for follow-up care.

Unfortunately, studying implicit categorization in humans is a challenge. Humans’ explicit and implicit cognitive processes can be deployed at the same time, with the final explicit decision influenced by both processes. The simultaneous processing by both systems has long been recognized in dual-system models of perceptual categorization (Ashby et al. 1998). In one experiment, Evans et al. (2019) intermixed images that contain only the implicit abnormality signal with images that contain an explicitly recognizable cancerous lesions. The result was a seeming loss of observer’s sensitivity to the implicit signal, though a post-hoc analysis identified that the most expert readers maintained their sensitivity to the implicit information. Clearly, the implicit processes are susceptible to being overridden or ignored in the presence of explicitly accessible features (cf., Kok et al. 2016; van Geel et al. 2017).

Most studies of medical abnormality detection rely on human observers; hence, these studies target the human visual system but struggle to isolate implicit visual categorization from explicit processes. However, non-human animals also possess highly capable biological vision systems. A frequent animal model in this line of “comparative visual cognition” studies is the pigeon, whose vision system shares some aspects with human visual systems (Cook et al. 2015). Neurologically, both primate and avian visual streams divide into two distinct pathways with stimulus processing more heavily processed in one of the two streams (Husband and Shimizu 2001). While the pairs of pathways are structurally homologous, the functional relationships are more complex. Humans rely on the cortical-thalamic pathway for our phenomenology of vision, while the collicular visual system is involved in orienting attention and some aspects of decision making related to visual search (Basso and May 2017; Tamietto et al. 2010). In pigeons, the inverse seems to be true, such that the structural homolog to collicular processing dominates their visual stimulus processing. Psychophysical investigations suggest that the resulting visual processing during search proceeds in fundamentally the same fashion in both pigeons and humans (Cook 2000; Cook et al. 1997; Cook et al., 2012a).

Studies of categorization, however, have been more equivocal in the comparison of human and avian visual processing. The pigeon (Columba livia), for example, has been shown to visually categorize a number of complex visual categories using an implicit process (Cook and Smith 2006; Qadri et al., 2014b; Watanabe et al. 1995). This categorization process seems devoid of various hallmarks of explicit cognitive processing (Qadri et al. 2019a; Smith et al. 2011, 2012), despite extensive training histories and long stimulus presentation durations. Some recent reports highlight how this independence from explicit human-like cognitive processing can benefit the pigeon if, for example, pigeons are faced with a complex stimulus space that would struggle under the explicit categorization mechanisms that humans report verbally (Wasserman et al. 2023).

Pigeons’ lack of explicit human-like visual cognition may be useful to isolate implicit processes in the tasks described above. Considering the problem of medical image readings – regardless of whether it’s a visual search task or a categorization task – pigeons may be an excellent candidate for a non-human biological model of visual implicit categorization. In support of this notion, Levenson et al. (2015) explored this possibility, training pigeons in a two-alternative choice task to report if a displayed image contained an abnormality. That investigation first examined standard histology slides (i.e., images of tissue examined under a microscope to reveal cell structures and properties), reporting that the pigeons both learned the abnormality discrimination and transferred it to novel images. This crucial test with novel, untrained stimuli identified that pigeons were not memorizing images but instead created a category representation of malignant histopathology useful for generalized discrimination. Another experiment in that same report examined masses in mammograms (i.e., X-ray images focused on breast tissue), but unlikely the histology, the pigeons only learned to memorize the images and could not generalize to novel image sets. Finally, a different paper (Navarro et al. 2020) examined pigeon’s reading of medical images enhanced by the standard colorization used to aid human visual processing and showed a considerable improvement in the detection of cardiac disease, suggesting that the examination of visual categorization of medical images in pigeons may confer practical or theoretical insight into humans’ processing of these images.

Here, we further consider the pigeon’s ability as an observer of a different imaging modality: the CT examination. Nodule detection in CT examinations is typically a complex extended activity across both space and time, distinct from the histology slides and mammograms reported by Levenson et al. (2015). CT scans use X-ray technology to produce a representation of the entire 3D object under study, which the radiologists then examine by looking through the 2D sections for features of interest. In detecting lung nodules, radiologists are trained to attend to the onset and offset of the nodule, distinguishing the nodule from vasculature which can look similar but ultimately attaches to other vasculature or anatomical structures in the CT scan. Previous studies have not used these types of dynamic medical imaging stimuli which require integration across space and time with pigeons. Other research has shown that pigeons have a remarkable sensitivity to motion-based features in visual discriminations (Cook et al. 2011), and are sensitive to the dynamic information available even when it is redundant with the static form information in the display (Koban and Cook 2009; Qadri and Cook 2017; Qadri et al., 2014b). Some of this dynamic processing differs from what we see in humans, with pigeons showing different responding to simple motion patterns (Hataji et al. 2019) or abstract motion patterns whereas humans perceive structured biological motion (e.g., Qadri et al., 2014a; Troje and Aust 2013). Critical for the current investigation, pigeons do possess the capacities to integrate and filter information that unfolds over extended presentations in ways that are consistent with our perceptions. For example, when presented with consistently occluded walking or running actions, pigeons could still learn the action categories despite never viewing the full displays (Gray et al. 2023). Even when some of the dynamic information is irrelevant or may even interfere with action categories, pigeons can still learn to attend to the category-relevant features to create broad categories (Cook et al. 2023). Thus, pigeons seemed a likely candidate for an animal system that could detect abnormalities in these CT scans.

In the current investigation, we conducted the first study of pigeons’ discrimination of dynamic medical images. We trained the pigeons to discriminate between healthy and unhealthy CT scans using solid lung nodules by presenting sequences of images. We evaluated whether any successful discrimination relied on simple memory or generalizable representations. Anticipating the need for generalization training given the results of previous reports (e.g., Levenson et al. 2015), we then planned to train the pigeons with additional stimuli. Thinking that we would find differences between animals trained to associate abnormality with food reward and animals trained to associate healthy displays with food reward, we tried to prime them towards a specific category to examine how attention to those specific features might affect their performance. Finally, we presented novel kinds of abnormalities to explore how generalizable a CT abnormality representation the pigeons were using.

Experiment 1

To explore pigeons’ utility as an animal model for radiologists’ visual implicit expertise, we trained pigeons to classify CT scans as “normal” or “abnormal” (i.e., containing a solid lung nodule) using a simple go/no-go discrimination paradigm. CT scans use X-ray imaging technology to create a detailed record of the structures inside of an organism. The resulting 3D information is then viewed by radiologists trained to “read” these scans by examining the 2D sections looking for signs of abnormalities. In looking for solid lung nodules, like those here, a radiologist might “scroll” through the images looking for the onset and offset of a bright region disconnected from other objects on the display. We considered that pigeons were likely capable of attending to this critical 2D feature of appearance or disappearance of the visual blobs. Alternatively, the pigeons could potentially extract other visual features to recognize the healthy or unhealthy scans, perhaps some prototypical nodule representation. In this latter case, we might anticipate finding a feature-positive effect (Sainsbury and Jenkins 1967), where pigeons’ ability to use the presence of a solid lung nodule as a defining feature could aid their discrimination. To examine this possibility, we considered whether the pigeons would demonstrate faster learning when the “abnormal” class was emphasized. We also considered the possible mechanism of discrimination being used, whether the pigeons were using a generalizable categorization process or simply memorizing the small set of stimuli we presented (cf., Levenson et al. 2015; Exp. 3 with breast masses). To evaluate which of these approaches the pigeons were using, we presented novel nodules from previously unseen patients along with novel healthy control stimuli.

Methods

Animals

Eight adult male pigeons (Columba livia) were used in the experiment, all 9 years old (except for two non-learning pigeons who were 18 and 21 years old). All pigeons had previous experience with motion discriminations in operant chambers, and they were maintained on a 12:12 L: D cycle at 80–85% of their free-feeding weight. The procedures used were approved by the Holy Cross Institutional Care and Use Committee.

Apparatus

Testing was conducted in a custom-built PVC (Type 3, gray) chamber (43 cm × 43 cm × 43 cm), equipped with an infrared touchscreen (ezscreen.com, EZ-170-WAVE-CW-USB) accessible through a 34 cm × 26 cm cutout behind which was placed a computer monitor (Asus VP228QG, 1920 × 1080 resolution at 75 Hz refresh rate). Below the touchscreen was a 2.5 cm × 2.5 cm cutout providing access to a grain hopper (Coulbourn Instruments) containing mixed grain. A houselight (Phidgets Inc., product ID 3609) centered in the ceiling continuously illuminated the chamber except during timeouts. All events were controlled by a Windows 10 PC using custom-written Python code built on PsychoPy (version 2021.1.4, Peirce et al. 2019). A USB relay board (Phidgets Inc., product ID 1017B) provided the interface for electronic circuits.

Stimuli

CT sections (slice thicknesses 1 to 2 mm) were obtained from a custom dataset of CT examinations of the lung and confirmed by clinical report, over-reading by 2 thoracic radiologists, and confirmation by an AI based nodule detection algorithm to either contain a lung nodule or have no lung nodules. These 3D volumes were displayed as orderly presentations of sequences of 2D sections. CT sections were loaded from jpeg files and displayed at 512 × 512 resolution, resulting in an image 12.7 cm × 12.7 cm on the display. The lung nodules varied in presence within the sequences, ranging from 4 CT sections containing nodules to 7 CT sections containing nodules (mean 5.375), though displays included one normal section on either end of the nodule-containing sections (watch Supplemental Fig. 1 or Supplemental Fig. 1 Annotated for an example). Presentation started on a randomly selected frame and progressed at a rate of 3 frames per second until the boundary normal section was reached and then it reversed progression until the other boundary normal section was displayed (at which point the progression reversed again and this process continued for the entire 10-s display interval). Using this presentation process, a nodule onset or offset could occur randomly throughout the entire 10-s display interval, preventing the pigeons from using simple time-related processing during the task. The number of onsets and offsets experienced within a trial depended on how many sections the nodules span (i.e., up to 6 onsets could be observed for nodules present across 4 sections, versus a maximum of 4 onsets for nodules present across 7 sections). For each nodule, a control stimulus with equal number of sections was selected from a patient without any lung nodules and with a similar overall appearance (as judged by a human experimenter; watch Supplemental Fig. 2 for an example). Initial training used 8 such nodules and their controls. During the novel item generalization test, 8 additional novel nodules and controls were tested using sections from previously untested patients and untested healthy controls.

Procedure

To evaluate if pigeons could detect solid lung nodules, we presented pigeons with 30 frames of CT scans displayed as a sequential “movie” over 10 s (see Fig. 1). Five pigeons were reinforced for pecking during a presentation containing an abnormality (Abnormal S+/Normal S- condition) while the remaining three pigeons were reinforced for pecking during “normal” displays (Normal S+/Abnormal S- condition).

Fig. 1

Sample stimuli used in these experiments, with expanded view of relevant structures at right. Each row depicts different display types, with each column showing consecutive sections. “Normal” and “Solid Lung Nodule” displays were used for training, with normal sections at the two ends. “Ground Glass Nodule” only has one healthy slice shown (for display purposes) and “Emphysema” has no healthy slices shown. For visual clarity, the rightmost column presents the abnormality (or normal vasculature) centered within an expanded display. To fully understand the visual presentation, readers are recommended to watch the Supplemental Videos

Pigeons were first trained to peck a “ready signal”, a white 3-cm circle positioned centrally in the display. Subsequently, after pecking the ready signal, a CT image was presented in the center of the display and updated every 333 ms. The pigeons were first trained to peck at all displays for a randomly timed food reward (variable interval VI-5 schedule). Once pigeons reliably pecked through the trial, they were started on the discrimination. During discrimination training, pecks continued to be rewarded on a VI-5 schedule for the reinforced S+ category, while for the non-reinforced S- category, pecks contributed to a dark timeout after the trial (1 s per peck). Discrimination training consisted of 96 trials per session, 48 from the reinforced category and 48 from the non-reinforced category randomly intermixed, conducted for 20 sessions. To evaluate peck rates to the reinforced category without the interruption of the food presentation, 8 S+ trials were programmed to withhold reinforcement.

After the initial training, a generalization test presented novel nodules to the pigeons who learned the discrimination, along with separate novel healthy control stimuli visually matched to the novel nodules. Such transfer trials were randomly interspersed within the session and provided neither reward nor timeouts. A total of eight novel nodules and 8 healthy controls were tested. These 16 untrained stimuli were randomly divided into two equal sets, with each test session presenting the stimuli from one of the two sets. Four total test sessions were conducted, so each novel stimulus was presented twice.

Discrimination performance was evaluated using Herrnstein’s rho, which is computed as a standardized Mann-Whitney U value (i.e., U observed divided by the maximum possible U). In cases where peck rate distributions to reinforced and non-reinforced stimulus categories (i.e., no discrimination) completely overlap, rho values would be 0.5, and in cases where the two distributions are complete separable (i.e. perfect discrimination), rho values would be 1.0. When considering novel transfer, only peck rates to novel stimuli contribute to the calculation.

Results

Six of eight pigeons learned the discrimination, with both non-learning pigeons in the Abnormal S+/Normal S- condition. Figure 2 displays this acquisition performance, with above-chance levels of discrimination by the end of 20 sessions of training (with non-learners t(7) = 3.9, p =.006, d = 1.4; graph and remaining statistics excludes non-learners, t(5) = 9.4, p <.001, d = 3.8; α = 0.05 for all tests). A mixed effects analysis of variance (ANOVA) with a within-subject effect of session block (five 4-session blocks) and a between-groups effect of stimulus assignment (Abnormal S + vs. Normal S+) confirmed only a main effect of session block (F(4,16) = 23.7, MSENum = 0.14, p <.001, η2p = 0.86), and no effect of stimulus assignment (F(1,4) = 0.2, p =.644) or its interaction with session (F(4,16) = 0.4, p =.839).

Fig. 2

Initial acquisition and performance with novel transfer stimuli, using average performance during each 4-session block of Experiment 1. Open symbols report baseline performance on the same sessions as the novel transfer tests were conducted. Error bars report standard error

All the pigeons who learned were then presented with eight novel nodules and eight novel normal displays, to evaluate if the discrimination would generalize to unseen exemplars. The pigeons successfully transferred to the novel exemplars (Fig. 2 “Novel transfer”), with average discrimination index of 0.715 which was well above chance, t(5) = 3.9, p =.006, d = 1.6. Despite its apparent reduction from the baseline training performance, a mixed effects ANOVA (within-subjects effects of novelty, between subjects effect of stimulus assignment) found no difference in performance between trained and untrained stimuli (F(1,4) = 4.7, p =.096). This is likely the result of variability between the birds, where two pigeons (one in each condition) had virtually no decrement (mean discrimination index for trained stimuli 0.828, untrained stimuli 0.824) while the remaining four demonstrated the considerable drop observed in the figure (trained stimuli 0.891, untrained stimuli 0.660). The ANOVA found no effect of stimulus assignment either on its own or as an interaction with the training/transfer condition (Fs(1,4) < 0.2, ps> 0.7).

Discussion

Generally, the pigeons were easily able to learn to discriminate these displays. Of the 8 pigeons trained, 6 learned to respond differently to displays where a nodule was present compared to when they were observing CTs without nodules. Crucially, these 6 pigeons did not simply memorize the displays, since their discrimination behavior transferred to novel CT sections containing novel nodules. This success suggests that pigeons may be an effective animal model for how a biological vision system might detect nodules in CT examinations. The two distinct stimulus assignment conditions appeared to make no difference to the pigeons’ performance. We anticipated that pigeons trained to peck to Abnormal displays might demonstrate a benefit if there were a specific feature included in the display, as would be expected for a feature-positive effect (Sainsbury and Jenkins 1967). However, no such effect was observed.

Two pigeons failed to learn the task. Both pigeons were elderly (18 and 21 years old), and both were in the Abnormal S+/Normal S- condition. Given that the three 9-year old pigeons in this group successfully learned, it seems likely that the non-learning is a consequence of the pigeons’ ages, though it is challenging to identify specific age-related causes for non-learning. A year after being removed from the study, the 18-year old pigeon was diagnosed with cataracts and ultimately removed from the colony altogether. The 21-year old, however, continued to appear healthy. There is little information about how aging affects pigeon cognitive abilities, though some memory changes have been noted (Coppola et al. 2015). Examining visual acuity from ages 2 y to 16 y, Hodos et al. (1991) noted a loss of visual acuity as measured using behavioral responses, despite subsequent physiological evaluation suggesting the eyes were capable of continued higher acuity. They conclude that some of the loss likely related to a decline in visual cognitive processing, though their older pigeons also had a different life experience that may have reduced their apparent acuity. Understanding what visual cognitive processing failures lead to elderly birds not learning this discrimination may separately reveal important visual processes in the analysis of dynamic medical images.

The pigeons’ transfer performance with novel stimuli was surprisingly not statistically distinct from their baseline training condition. This was surprising because the pigeons showed a reasonable reduction in performance, comparable to most tasks with stimulus sets of this size and these many animals (Cook et al. 2015). As discussed above, this non-significant result likely stems from a mixture of pigeons who did not demonstrate a reduction of performance from pigeons who did have reduced performance with the novel stimuli. The typical reduction of discrimination during generalization testing with pigeons has generally had two competing interpretations (e.g., Levenson et al. 2015). In both, the reduction is taken to indicate that the pigeons had been partially memorizing the stimuli. However, a generous interpretation of that memory use is that it performs akin to familiarity, not as a central part of the discrimination, and since the stimulus is unfamiliar it has less association with the reinforcement contingencies. In this interpretation, the performance decrement indicates that the pigeons recognize that the stimuli were in fact novel and therefore perceptually distinct from the training stimuli. This decrement becomes a necessary feature to assert that the pigeons’ generalization reflects visual categorization mechanisms and is not a simple byproduct of perceptual confusions. A less generous interpretation suggests that the pigeons’ algorithm partially relies on an exact match to memorized exemplars. In this latter case, further training and testing is typically employed to reduce the pigeons’ reliance on memory processes when being presented with a small training set of stimuli. We address this possibility in the next experiment by increasing the number of stimuli during training.

Experiment 2

Since the pigeons successfully learned to detect nodules in the CT scans, we turned to increasing the stability of the discrimination by adding more stimuli into the training set. For this, we first decided to include the stimuli used to evaluate transfer in Experiment 1 into the training set of stimuli. We also used visual transformations (mirror images) to increase the number of apparent distinct stimuli, a frequent measure used to increase the discrimination training set in both animal experiments and computer vision training designs (e.g., Levenson et al. 2015; Shorten and Khoshgoftaar 2019). By incorporating simple manipulations, the training classifier (pigeon or computer) becomes less reliant on simple low-level cues available in the general pattern of the classes. For example, if the training sets coincidentally had more nodules in the left lung than the right lung, a classifier might become attuned to that spatial disparity. By training the classifier on a stimulus and its mirror image, then both lungs are made equally likely to be relevant to the discriminative task at hand. We again suspected that a difference may appear between reinforcement conditions as the task becomes more challenging, with more stimuli employed in each class. The presence of a nodule continued to seem like a concrete feature that could promote ready discrimination. Thus, after training with the additional stimuli, another generalization test was conducted to evaluate how the pigeons were categorizing the displays and to look for potential group differences based on the reinforcement assignment.

Methods

Animals, apparatus, and stimuli

The same animals and apparatus were used as at the end of the previous experiment (i.e., 3 pigeons in the Abnormal S+/Normal S- condition and 3 pigeons in the Normal S+/Abnormal S- condition). The same stimuli from the previous experiment were used, along with their mirror images. For novel transfer, a new set of 8 nodules and 8 healthy controls were used. Because of the limited number of patients available, some nodules originated from previously used patients but presented novel CT sections with new nodules not previously seen from these patients (4 of 8 nodules were from previously used patients; 1 of 8 healthy controls).

Procedure

Immediately after the generalization test of the previous experiment, the pigeons were presented with sessions that incorporated the expanded training stimulus set. In addition to the original 16 stimuli, the novel transfer stimuli were included for a total of 32 distinct stimuli (16 Abnormal, 16 Normal). To increase the apparent number of stimuli, these 32 stimuli were mirrored about the body’s midline to produce 32 additional stimuli. Only a mirror image transformation was used to increase the number of stimuli since other manipulations (such as rotations) would not align with the typical presentation style used to evaluate these stimuli. Sessions continued to consist of 96 trials, 48 from the pigeons’ S+ class and 48 from the pigeons’ S- class. Thus, each trial presented a randomly selected stimulus from the set of 64 (16 Abnormal, 16 Normal, 16 Abnormal mirrored, 16 Normal mirrored) ensuring that equal numbers of Normal and Abnormal stimuli were presented. To measure peck rates of the S+ class, 8 S+ trials were programmed to withhold reinforcement. As in Experiment 1, a total of 20 sessions of training were conducted. After completing the generalization training, the pigeons were again evaluated for transfer across four test sessions in the same procedure as the previous experiment except the second presentation of each stimulus now presented the mirror image version of the novel stimulus.

Results

The pigeons continued to discriminate the classes of stimuli effectively, and they demonstrated little change in discrimination over the course of the 20 sessions (Fig. 3). Because of the randomness in the likelihood of items being selected for each session, the discrimination index was computed for each stimulus subset within each four-session block. These data were then examined by a mixed effects ANOVA (within subject factors: stimulus set, mirror image, session block; between groups factor: stimulus assignment). This analysis confirmed no main effect of session (F(4,16) = 0.4, p =.826), nor a main effect of the mirror image or stimulus assignment (Fs(1,4) < 0.2, ps > 0.679), but it did identify a main effect of stimulus set (F(1,4) = 12.9, p =.023, MSENum = 0.084, η2p = 0.76) originating from better performance for stimuli from the original training set. The analysis suggested a possible interaction between stimulus set and session (F(4,16) = 3.1, p =.047, MSENum = 0.011, η2p = 0.44), where the main effect of stimulus set was reduced during some of the interior sessions. No other interactions with session were significant Fs(4,16) < 1.8, ps > 0.18. Thus, there was no systematic improvement during the course of the training. However, in addition to the main effect of stimulus set, an interaction between stimulus set and mirror image presentation use (F(1,4) = 12.5, p =.024, MSENum = 0.010, η2p = 0.76) was found. This interaction reflects how the pigeons’ discrimination with the original training set was better for un-mirrored stimuli, while for the newer stimuli, the mirror image version supported better discrimination, likely reflecting differences in the stimulus sets for which lung likely had the nodule. A significant three-way interaction between training set, mirror image, and stimulus assignment (F(1,4) = 14.9, p =.018, MSENum = 0.011, η2p = 0.79) suggests that the difference on mirrored versions of the stimuli was a stronger effect for the pigeons in the Normal S+/Abnormal S- stimulus assignment compared to the Abnormal S+/Normal S- birds. All remaining interactions were non-significant (Fs < 4.3, ps > 0.107; see Supplemental Fig. 3).

Fig. 3

Discrimination as a function of addition training with new or mirrored stimuli, using average performance during each 4-session block of Experiment 2. Sets here refer to the original training set (Set 1) or the newly added stimuli (Set 2)

After incorporating these transfer stimuli (and their flipped versions) into training, testing of another novel eight nodules and 8 novel controls replicated the results from Experiment 1. The pigeons’ average discrimination index (M = 0.64, SE = 0.019) was above chance (t(5) = 7.3, p <.001, d = 3.0; no between-group difference t(4) = 0.4, p =.691), but it was not much different from Experiment 1’s transfer performance (t(5) = 1.6, p =.170).

Given the stability of the discrimination during this training, we used this period of 20 sessions to evaluate the pigeons’ pecking behavior to better understand their processing of the stimuli during the display. Figure 4 shows one such analysis, examining how the pigeons’ perceptions of the display changed throughout the course of the trial. This figure uses a normalized peck rate to account for variation across the pigeons’ peck rates by dividing the birds’ peck rates by their peck rate on the last 2 s of the display on S+ trials (i.e., \(\:Normalized\:Peck\:Rate=\:\frac{Peck\:Rate}{Terminal\:S+\:Peck\:Rate}\). The time blocks used in the figure (333 ms) equates with the frame presentation duration, so that each data point identifies peck rates during a single frame’s duration during a trial. As Fig. 4 shows, after pecking the ready signal to start the trial, pigeons waited to observe the stimulus before systematically and regularly responding. Pecking during the first frame of the trial shows no evidence of discrimination, but by the second frame, peck rates increased overall and the S + and S- categories started to separate. Clearly, the pigeons did not require the full 10 s to decide on whether a given trial contained an abnormality, as the peck rates reach an asymptotic level by about 3 s. Curiously, regardless of the pigeons’ stimulus assignment, the normal stimuli (circles) in the figure reach their asymptotic peck rates after 1 s while the stimuli containing nodules (triangles) require an additional 1 s before reaching asymptote.

Fig. 4

Peck rates during a trial during Experiment 2. Solid symbols represent pecking during the assigned S+ stimuli (paired with reinforcement) while open symbols represent pecking during the assigned S- stimuli. Each data point reports responding during the 333 ms of a single frame. Triangles represent pecking during stimuli with a nodule and circles represent pecking to the healthy control stimuli

Next we tried to evaluate if there were regularities in how the pigeons were processing the displays. Specifically, we examined whether the pigeons’ performance suggested an identifiable “ranking” of stimuli from easy to hard. If some stimuli were easily classified while others were more challenging, we could begin to develop a prototypical representation of the two classes of stimuli. To this end, we considered each group of birds separately, hypothesizing that the pigeons assigned to peck to Abnormal stimuli (Abnormal S+/Normal S-) might use distinct properties from those assigned to peck to Normal stimuli (Normal S+/Abnormal S-). For each stimulus class, we computed the Z scores within each pigeons’ peck rates. We then used a simple linear model with a fixed effect of stimulus ID and a random effect of bird to see if a significant amount of the variance could be accounted for by the stimulus identity. This analysis suggested that the abnormal stimuli were more consistently responded to overall, suggesting a more consistent set of identifiable features than normal stimuli. The Normal S+/Abnormal S- birds seemed to respond to abnormal displays in a reliable fashion F(15,30) = 4.5, p <.001, and Abnormal S+/Normal S- birds approached significance but did not reach it F(15,30) = 1.9, p =.066. Thus, as might be expected, the distinct nodule feature produced more systematic processing and responding than the more varied Normal category (Fs(15,30) < 1.5, ps > 0.161).

Discussion

The addition of more stimuli into each of the stimulus classes did not seem to change the pigeons’ performance. Indeed, the performance curves generally seemed flat for the duration of the training phase. This flatness suggests that the pigeons had learned an effective, generalizable discrimination from the initial 8 exemplars of each category. We then examined the pigeons’ peck rates to try to gain further insight into the cognitive processes being applied. The class of Normal stimuli produced a clear response within 1 s, while Abnormal stimuli required an additional 1 s of processing. Peck rates also revealed how the Abnormal stimuli seemed to produce more systematic responding, as seen in which stimuli were “easier” across the birds, while Normal stimuli did not. This seemed clearer in the “avoid” or “no-go” class of stimuli than the “go” class.

The analysis of the pigeons’ peck rates raises two interesting questions. First, what is the process by which the stimuli are being evaluated on a trial-by-trial basis? Presumably, the pigeons process the displays in a single, systematic fashion. How then do the two classes of stimuli produce different processing curves during the trial? The stability of Normal stimulus responding after 1 s suggests that these are systematic and easier to discriminate compared to the Abnormal stimuli. However, the pigeons show more agreement about the Abnormal than the Normal stimuli. Perhaps the issue is that a Normal signal is somewhat diffuse as healthy tissue has high variability, and healthy tissue is also present in the healthy part of the lung of patients with nodules. In this case, on trials with a nodule, the pigeons might continue scanning the display and shift their decision partway through the trial. This hypothesis does not account for the clear separation of peck rates in the first second of the trial. An alternative mechanism involves systematically evaluating the stimuli through a sequence of features. In this alternative, perhaps the features evaluated early may be more diagnostic of healthy stimuli compared to the ones evaluated later. These early features might then be more similar to those used in human implicit processing, while the later features might match those that support explicit processing. This is discussed further in the General Discussion.

A second interesting question regards the pigeons’ agreement in responding to Abnormal stimuli. One interpretation of this pattern is that some of the stimuli are “easy” examples (category-typical class members) or while others are “hard” (less categorically typical). If this is the case, the pigeons’ relative difficulty judgment does not match the easy/hard judgments provided by two of the expert radiologists who reviewed the stimuli prior to their use. Radiologists rated the stimuli on a scale of 1–4 prior to the start of the experiment, and we had selected from the stimuli in a counterbalanced fashion. Using these difficulty values and the pigeons’ average response rankings shows little relationship (using individual birds’ values rs = − 0.13, − 0.10, and 0.02), suggesting that the pigeons’ and humans’ feature sets for this discrimination differ, at least when comparing to radiologists’ explicit difficulty rating of each nodule. Nevertheless, systematic responding suggests that pigeons are attending to common features in the displays.

Experiment 3

One goal of the current study was to identify possible differences in the cognitive processes used to categorize displays as Normal or Abnormal. Asymmetries in the approach to a categorization task can result from differences in the frequency and distribution of the different categories (i.e., “Normal” and “Abnormal” categories; Kruschke 2008). Asymmetries are especially likely if using a simple template-matching paradigm to identify visual categories (Sinha and Balas 2008) when those categories are potentially quite variable. For the current task, for example, there are likely many variations of Normal to learn about as healthy tissue has wide variability while the characteristics of a “nodule” are likely more limited to the features around the shape and onset/offset of the nodule. To this end, we have repeatedly evaluated group differences based on stimulus assignment, but no clear statistically significant differences have been observed.

In this experiment, we attempted to accentuate possible differences by priming the pigeons towards their S- category. This manipulation had two goals. First, if we believe that the two tasks the pigeons were given are asymmetric, then this priming should produce a more dramatic difference. We had approached this investigation believing that the pigeons would show a similar behavior in our discrimination task as they have in previous comparable reports (Cook et al. 2012b; Qadri et al. 2019b; Qadri et al., 2014b). In these designs where stimuli are presented for an extended period of time and responses may trigger reinforcement throughout the trial, pigeons have demonstrated high peck rates at the start of the trial and on trials where S- stimuli are recognized, they stop pecking, typically for the duration of the trial. In cases where information changes throughout the trial and perceptions can change as a result (Cook et al. 2015), pigeons resume pecking and potentially stop again; however, the majority of the discrimination seems to result from the initial “stop-pecking” decisions. Given this context, we increased the number of S- stimuli we presented. Second, we believed the increased prevalence of the S- stimuli could cause the pigeons to adopt a specific strategy related to the S- stimulus class, as is seen in cases of base-rate priming and over selection during visual search (Fremouw et al. 1998; Langley et al. 1996). In the current design, we would expect a type of feature-positive effect, where the pigeons in the Normal S+/Abnormal S- category would ultimately perform better because of the useful features in the S- class. This priming training should therefore increase the systematicity in responding across birds.

Methods

Animals, apparatus, and stimuli

The same animals and apparatus were used as in the previous experiments. The stimuli from Experiment 2 (including mirrored versions) were re-used, and a novel set of stimuli were introduced in the generalization test at the end of the training sequence. These novel stimuli were from previously presented patients, but new Abnormal stimuli were created from nodules not previously seen before and new Normal sections were used.

Procedure

The pigeons were primed to expect their S- category by increasing their relative frequency from 50% to 75% of the trials during training. While the session still consisted of 96 trials, 72 of these were now S- trials, and 24 were S+ trials. For the pigeons in the Abnormal S+/Normal S- assignment, this meant 72 Normal stimuli and 24 Abnormal stimuli were presented, while for the pigeons in the Normal S+/Abnormal S- assignment, this meant 72 Abnormal stimuli and 24 Normal stimuli were presented. During this training, stimuli were presented as flipped 50% of the time to maintain the increased variability in stimulus sets from the previous experiment. For S- stimuli, 64 of the 72 trials presented each stimulus (and its mirrored version) twice, along with 8 randomly selected stimuli that were presented an additional time. The S+ trials still had 8 trials programmed to provide no reinforcement to allow us to measure peck rates without the interruption of the hopper. After 25 sessions of this weighted training, the pigeons were tested with six novel nodules and six novel controls. Six nodules were used instead of eight because of limitations on the number of nodules available for testing. These test sessions consisted of the same 96-trial baseline (with the additional S- stimuli) and 6 additional trials testing 3 novel nodules and their control stimuli. Four test sessions were conducted with at least one baseline session between test sessions, following the same procedure as Experiment 2.

Results

As with the generalization training, the pigeons’ discrimination performance did not appear to change during the extended training period. Figure 5 shows this fairly stable performance pattern, though it suggests that the Abnormal S+/Normal S- condition might have improved a little over the course of training. A mixed effects ANOVA (within subject: session, between groups: stimulus assignment) detected no significant effect of session or its interaction with stimulus assignment (Fs(4,16) < 1.3, ps > 0.331). The “Transfer” portion of Fig. 5 shows that the birds continued to generalize with above-chance performance as a whole, t(5) = 3.9, p =.011, d = 1.6, but there remained no significant difference between the groups t(4) = 1.3, p =.258 and no significant change within each group compared to the prior generalization test ts(2) < 2.3, ps > 0.145.

Fig. 5

Discrimination as a function of additional training with base-rate priming emphasizing S- stimuli, with subsequent discrimination transfer to novel stimuli using average performance during each 4-session block of Experiment 3. Open symbols report baseline performance on the same sessions as the novel transfer tests were conducted. Error bars report standard error

Another change that the priming could have produced is more systematic evaluation of the S- classes. Thus, we re-evaluated the peck rates to the stimuli to see if the pigeons’ responses were again more or less in agreement about which were the easier stimuli. This replicated the results in the previous section, where only the pigeons in the Normal S+/Abnormal S- had agreement about their peck rates to S- stimuli (F(15,30) = 2.7, p =.010).

Discussion

Here we attempted to make potential group difference more dramatic by priming the groups to their S- conditions. The discrimination seemed unchanged, as measured by ongoing discriminative performance. In examining transfer to novel stimuli, no between-group differences emerged. Whether this reflects that there truly is no difference in whether the birds are attending to the nodules vs. healthy tissue or if it instead reflects an underpowered comparison is unclear. Using the effect size seen in this study to estimate a design that would have sufficient statistical power suggests that a study with 24 pigeons (12 in each group; analysis conducted via G*Power) would be able to detect this size of an effect 80% of the time. The lack of a significant difference between the groups suggests that either we have insufficient power to detect a difference or that this is a paradigm that does not produce the typical feature-positive effect.

The data show that pigeons in the Normal S+/Abnormal S- group have comparable relative peck rates to the different Abnormal stimuli. In a fashion, this suggests that the pigeons are sensitive to the same Abnormality-specific characteristics when they engage in “stop pecking” decisions. The lack of a comparable significant relationship across “stop pecking” decisions to Normal stimuli suggests that the pigeons’ processing of Normal stimuli is less consistent, perhaps because the “Normal” is more variable and lacks a specific “healthy” signal.

Experiment 4

After the initial training with 8 stimuli of each type, the pigeons showed continued effective generalization to 22 new stimuli of each class (8 transfer stimuli in Experiment 1, 8 transfer stimuli in Experiment 2, and 6 transfer stimuli in Experiment 3). This kind of generalization is generally not seen with pigeons, who may require tens of stimuli before being able to effectively categorize novel stimuli (Daniel et al. 2015). Given the effective generalization of the current discrimination, we examined the breadth of the pigeons’ generalization by presenting perceptually novel abnormalities (refer to Fig. 1, right). We examined two particular kinds of untrained abnormalities. In the “ground glass nodule” type of abnormality, the nodules did not have well-defined edges, instead appearing like “ground glass” in the CT scan. In the “emphysema” type of abnormality, the visual signature was the opposite of what the training solid lung nodules appeared like – dark absences instead of bright presences. Further, the “emphysema” scans do not contain healthy tissue bordering the abnormality, thus making this transfer even more difficult. In this way, the ground glass nodule acts as a kind of near-transfer, and the emphysema acts as a kind of far-transfer.

Methods

Animals, apparatus, and stimuli

The same animals and apparatus were used as in the previous experiment. The training stimuli and conditions were the same, but new abnormalities (“ground glass nodules” and emphysema; see Fig. 1) were shown during the transfer test. These abnormalities are visually distinctive compared to the solid pulmonary lung nodules the pigeons received during training. “Ground glass” nodules have a semi-transparent appearance, compared to solid lung nodules, and poorly defined edges (watch Supplemental Fig. 4 or Supplemental Fig. 4 Annotated for an example). Emphysema presents as dark regions in the CT image, an inversion of the relevant contrast relationship when detecting solid lung nodules (watch Supplemental Fig. 5 for an example). Emphysema is present throughout the lung, so the abnormality is present in every slice in the movie. Hence, the emphysema stimuli do not have a bordering normal slice of healthy tissue above and below the abnormal tissue. As before, untrained Normal control images were used as comparisons to the novel Abnormal images, and while these originated from patients previously seen, we tested new sections of the lung that the pigeons had not seen previously.

Procedure

Two weeks after the completion of Experiment 3, during which time the pigeons continued to get the S- weighted training sessions described in 4.1.2, the pigeons were tested with sessions that presented the visually novel abnormalities. Each test session contained two novel abnormalities of each type along with the novel Normal stimuli (though, as in the previous experiment, these Normal sections originated from patients previously presented to the pigeons, but these sections of the lung had not previously been seen). In two separate test sessions, these test stimuli were presented as random non-reinforced insertions into the baseline sequence of trials. Test sessions were thus composed of 96 baseline trials (72 S+, 24 S- trials) plus two additional trials each of “Ground glass” nodules, their controls, emphysema, and their controls (2 × 4 = 8 transfer trials) for a total of 104 trials. Two additional test sessions were conducted that re-tested these with mirror image transformations of the novel stimuli. Finally, this sequence of four sessions was repeated so that this experiment presented comparable counts of novel stimulus presentations as the previous experiments (i.e., 8 total test sessions presenting each abnormality 16 total times). To minimize learning, test sessions were separated by at least one baseline session.

Results

During test sessions, the pigeons’ performance remained excellent with trained items and showed considerable transfer to the novel ground glass nodules and emphysema. Figure 6 depicts performance during this phase, separated by each type of abnormality. Pigeons’ discrimination index remained above-chance for both new abnormality types (ts(5) > 4.6, ps < 0.006, ds > 1.9), though there was a clear decrement in performance for the novel ground glass nodules and the novel emphysema displays. This latter effect was confirmed with a mixed ANOVA (within subject: abnormality type, between groups: stimulus assignment) which identified a significant main effect of abnormality type F(2,8) = 20.5, MSENum = 0.11, p =.001, η2p = 0.84. Uncorrected post-hoc pairwise comparisons identify that the trained solid lung nodule performance was different from each of the novel abnormalities (ps < 0.011) while the two novel abnormalities were not significantly different from each other (p =.316). Stimulus assignment continued to have no significant main effect (F(1,4) = 0.3, p =.595) nor interaction with stimulus type(F(2,8) = 0.05, p =.949).

Fig. 6

Transfer performance to novel types of abnormalities during Experiment 4. SLN = solid lung nodules, GG = ground glass nodules, and EM = emphysema. Error bars report standard error

Discussion

Without any further training, the pigeons’ trained discrimination with solid lung nodules produced significant transfer to novel visually distinct abnormalities that differed considerably from the original training stimuli. Unlike the well-defined temporal and spatial edges of solid lung nodules, ground glass nodules fade towards their edges and appear translucent on a CT examination. Emphysema presents a different temporal signature because it tends to be present throughout the lung tissue, and it presents the opposite contrast signature (i.e. dark spots). Nevertheless, the pigeons responded to these novel and distinct visual features consistent with their training; they were capable of implicit categorization of these visually novel abnormalities.

A critical analysis of our design might suggest that the pigeons’ successful transfer is an artifact of our methodology. The Normal stimuli used in our transfer test in this experiment used novel CT sections that the pigeons had not been trained on or seen before, but these sections originated from individuals that the pigeons had previously seen or been trained on. This previous exposure raises concerns that idiosyncratic aspects of the lung tissue (or even content from outside the lung tissue) could support systematic responding to the Normal stimuli. Because our response metric reports separability, this could then support discrimination if the Normal controls were identified using individual-recognition processes and not a perceptual categorization process related to the abnormalities, assuming that the novel Abnormal stimuli produced intermediate responses. This assumption does not bear out, with the Abnormal S+/Normal S- birds showing fairly consistent peck rates between the trained solid lung nodules and the novel abnormalities. To be clear, pigeons responded similarly to both the highly trained abnormalities as well as the novel visual abnormalities. To fully address this possibility, a subsequent study would need to be conducted that presented all-novel stimuli to ensure that individual familiarity is not a contributing source of discrimination.

Pigeons trained to peck at abnormalities pecked more to displays with these novel abnormalities and pigeons trained to avoid pecking at abnormalities pecked less to these displays than to corresponding novel healthy stimuli. This successful transfer suggests that the pigeons were not simply responding to the superficially obvious and explicitly reported visual appearance and disappearance feature (or change in the display) used by radiologists for solid lung nodules. The stark appearance/disappearance feature is not shared with ground-glass nodules and particularly with emphysema where the abnormality is spread across every slice of the CT. Instead, the pigeons are likely attending to some larger configural feature that is not well-analyzed or understood in these stimuli. What this feature might be is discussed in the General Discussion below.

General discussion

Six pigeons learned to discriminate CT scans containing a solid lung nodule from normal displays (c.f., Levenson et al. 2015) and did so using a generalizable mechanism. Training the pigeons with additional novel stimuli and priming them towards a particular category did not seem to substantially change their performance. The category learning then transferred without specific training to visually distinctive, new types of abnormalities (“ground glass nodules” and emphysema). These experiments reveal that the complex, attentionally demanding activity of evaluating CT examinations for nodules can be achieved by a biological vision system like the pigeons’, which seems to only use implicit categorization mechanisms. Additionally, the successful transfer of discrimination from solid lung nodules to ground glass nodules and emphysema without further training suggests that these three visually distinct abnormality types may have a common perceptual signature that the pigeons were able to detect implicitly.

Detecting abnormalities in CT scans is a considerably more complex visual discrimination than most investigations using the pigeon model. In the detection task here, the target nodule is a localized region that radiologists report as looking similar to the surrounding vasculature at any given moment. Only through extended observation of the nodule region across space and time is the nodule identifiable, but nothing visually identifies the region for the observer. Thus, the pigeons needed to examine the entire display to identify a region in which a nodule is appearing and disappearing over time while ignoring the irrelevant variation that presumably captures their attention. This degree of dynamic information processing has not been observed in pigeons before. Previous reports of temporally constructed information, such as in completion of occluded action (Gray et al. 2023), only required the combination of the dynamic display features presented. Here, the display includes irrelevant dynamic information that would make a simple dynamic-relevance heuristic more challenging. One previous study with pigeon presented irrelevant camera motion while presenting an action category (Cook et al. 2023), which could be similar to the discrimination here, but in that case secondary form features were present that could aid the discrimination. Thus, this report extends the excellent visual processing abilities of pigeons into a new territory, one defined by constructed dynamic features capable of categorization in the presence of irrelevant dynamic noise.

This interpretation, however, assumes that the pigeons are using the same visual feature as humans to identify the nodules. This assumption is challenged by known differences in how pigeons approach visual discriminations compared to humans. In many investigations, pigeons have repeated shown a propensity to attend to spatially narrow or local details while humans generally proceed with a globally biased visual process (Cavoto and Cook 2001; Murphy et al. 2015; Troje and Aust 2013). While pigeons and humans are capable of shifting their attention to the relevant features as needed (Fremouw et al. 2002; Lea et al. 2013, 2018), differences in spatial precedence make it unlikely that the same features are used by both species. The successful transfer of discrimination to the novel abnormality types in the current investigation is also a challenge on this front. The trained solid lung nodules share few visual features on a CT examination with ground glass nodules and emphysema.

Explaining the pigeons’ successful discrimination with these displays might then suggest that some shared feature exists between these displays. Possibly relevant, the time course of pigeons’ peck rates revealed a change in the stimulus processing after approximately 1 s of viewing. Perhaps discrimination prior to this 1-s interval used abnormality-shared features and discrimination after that interval used abnormality-unique features. One possibility for this shared feature is some type of globalized signal in the scan indicating that the nearby tissue is unhealthy. Pigeons have previously been shown to process such global features particularly when the local features are densely packed (Goto et al. 2004) as is the case in a CT of the lung. This global “unhealthiness signal” has been suggested to facilitate the processing of implicit information in radiologists, permitting the detection of related abnormalities while observing otherwise healthy tissue (Evans et al. 2016, 2019). A contrasting theory suggests that a “normal template” facilitates these implicit discriminations in humans (DiGirolamo et al. 2023). Both mechanisms could have facilitated the transfer of discrimination in comparable ways, so future research will have to examine whether these abnormalities share features using methods focused on that task. Tasks using stimulus manipulation can start to identify what features the pigeons are attending to. Analysis of local/global feature use can be accomplished through spatial frequency filtering, for example, and may reveal that pigeons’ discrimination either uses the small, detailed features (high spatial frequencies) or separated visual features (low spatial frequencies). A temporal analysis would also be interesting, by either presenting single frames for the entire display duration or presenting frames in an inconsistent order. Successful discrimination in this latter case, though impossible from a human radiologists’ explicit processing of the displays, might reveal that non-temporal cues are present in the displays that the pigeons may be able to use to categorize a CT section. Whether this is possible remains to be evaluated empirically.

One goal for the current study was to examine whether the pigeons’ discrimination would benefit from a focus on the “normal template” by priming one group of birds to focus on features related to that category while the other group was primed towards the abnormalities. While a statistically significant difference between the two priming groups was not observed, the birds primed to normal stimuli may have mildly benefited from their S- priming. This hinting comes from small non-significant increase in discrimination ability when primed to attend to “normal” stimuli with concordant less decrement with novel stimuli (see Fig. 5). If the use of a normal template were more effective than employing search for the more-variable abnormality targets, this would have considerable impact on radiologists’ training. For example, pathology examination may be recommended to follow a structured hierarchical decision process to first determine non-normalcy and then advance to localization and identification of pathology. Thus, another goal for future research would be to re-examine this specific question with a higher-powered design.

The detection and classification of abnormalities is also an active research area for artificial intelligence (i.e., computational vision). Traditional machine learning systems, designed in conjunction with experts to identify meaningful features that correlate with the categories of interest, have been less successful than more recent “deep learning” networks that are less scrutable (for reviews, see Gu et al. 2021; Huang et al. 2017). Likely, some errors by the traditional systems relate to the bias towards explicit mechanisms prevalent in human explanations of detecting the nodule. While relevant to how humans describe and discuss the abnormalities, such explicit features may have less utility than simpler perceptual features. Pigeons seem to lack the explicit categorization mechanisms used by humans during simple perceptual learning tasks (Smith et al. 2012). These birds may potentially serve as an animal model of biological vision, and like deep learning neural networks, they lack humans’ explicit bias. Given the birds’ success here, some relatively “simple” featural analyses (i.e., by a deep neural network) could be developed to successfully identify abnormalities in CT scans. Critically, the pigeons’ transfer to multiple visually distinct abnormalities suggests that there is a critical feature shared across these visually distinct abnormalities. Hence, a single machine learning algorithm could be developed to successfully detect multiple abnormality classes instead of the abnormality-specific detection algorithms that currently prevail (Milam and Koo 2023). Further studies of this animal model have the potential to determine these shared perceptual characteristics, which may enable machine learning systems to produce generalizable abnormality detection algorithms.

Investigations like these may ultimately suggest changes to the training of future radiologists. This would require careful work during translation since identical manipulations could yield conflicting results. For example, the method of priming used here has demonstrated differences in amount of processing of primed and unprimed information in pigeons (Fremouw et al. 1998), and do not produce comparable response biases that the same procedure produces in humans (Hartl and Fantino 1996). Instead, the application of these results needs to be applied with sensitivity to each species’ specific processing streams. In particular, these data suggest radiologists’ education may benefit from focusing on visual perceptual learning and execution (see DiGirolamo et al. 2023) instead of focusing on didactic instruction and explicit feature reporting. As long as expert radiologists are examining medical images in their standard practice, engaging in this same dynamic visual categorization task, further developing and understanding a biological model of this visual cognitive process remains critical.

References

Download references