Skip to main content
Have a personal or library account? Click to login
Cognitive Human Factors and Forensic Document Examiner Methods and Procedures Cover

Cognitive Human Factors and Forensic Document Examiner Methods and Procedures

Open Access
|Jun 2020

Full Article

Introduction

The extensive scrutiny of the methods and findings of numerous areas of expert testimony has prompted acrimonious debate among academicians, forensic practitioners, and legal professionals concerning what has been referred to by the Forensic Science Committee of the National Academy of Sciences (“Committee”) as “faulty forensic science analyses” [1, p. 4]. While acknowledging the importance and utility of the forensic disciplines, the Committee also addressed the perceived flaws in such evidence. For example, advances in technology in various forensic disciplines, especially in the field of DNA testing, show that erroneous or misleading forensic evidence has contributed to the wrongful conviction of innocent individuals [1]. The Committee called for improvements in forensic science practices, arguing that increased and demonstrated reliability and validity in forensics will help law enforcement investigations by improving the reliability of identifications. Additionally, homeland security efforts will also improve as advancements are made in the methods and procedures of the forensic disciplines [1].

The Committee specifically identified several important issues, including practitioner certification, accreditation, and the availability of skilled, well-trained personnel [1]. Many areas of forensic science lack of uniformity in training, accredita tion, and practice standards. The report stated that operational principles and procedures for many disciplines are not standardized between or within jurisdictions, attempts at standardization are not viewed favorably in many instances, and that even protocols such as Scientific Working Group standards “often are vague and not enforced in any meaningful way…These shortcomings obviously pose a continuing and serious threat to the quality and credibility of forensic science practice” [1, p. 6].

The Committee also discussed the lack of demonstrated validity and reliability within the interpretation-based disciplines, stating “… no forensic method has been rigorously shown to have the capacity to consistently, and with a high degree of certainty, demonstrate a connection between evidence and a specific individual or source…The simple reality is that the interpretation of forensic evidence is not always based on scientific studies to determine its validity. This is a serious problem. Although research has been done in some disciplines, there is a notable dearth of peer-reviewed, published studies establishing the scientific bases and validity of many forensic methods” [emphasis added] [1, p. 7-8].

The 2016 report of the President’s Council of Advisors .on Science and Technology titled Report to the President on Forensic Science in Criminal Courts: Ensuring Scientific Validity of Feature-Comparison Methods1 similarly identified what the Council concluded were two important research gaps: A need for clarity about the standards for reliability and validity in the forensic sciences, and the need for ongoing studies to establish the extent of the validity and reliability of the forensic sciences.

The challenges of implementing common standards for reliability and validity across forensic fields have been illustrated by the ongoing debates about which agencies or organizations should make these determinations. According to Dror, empirical observation is neither strictly objective nor subjective, but instead exists on a continuum along which most forensic inquiries fall [2]. Consequently, creating a dichotomy for fields of inquiry based on the extremes of pure objectivity and pure opinion is neither accurate nor fruitful. Dror stated that even subjective, experience-based expert opinion can be accurate and valuable, despite the possibility that such expert opinion may be at increased risk for error, bias, and contextual influences. He wrote “even with quantification and statistical tools, the human element still plays a critical role, and therefore cognitive issues may continue to play an important role even in the less subjective domains of forensic science” [2, p. 81].

Current Approaches to Understanding the Examination Process

Identifying and mitigating factors that may contribute to misleading conclusions has long been a matter of concern for many forensic fields. Forensic science is produced and consumed in the context of various systems developed by human actors. Recent approaches to understanding the impact of potentially biasing information on forensic conclusions have embraced a human factors perspective. The International Ergonomics Association defines human factors as “the scientific discipline concerned with the understanding of interactions among humans and other elements of a system, and the profession that applies theory, principles, data and methods to design in order to optimize human well-being and overall system performance” [3]. Cognitive ergonomics is a subfield of human factors in which cognitive processes such as memory, perception, reasoning, decision making, skilled performance, and human reliability are examined in the context of work and operational settings. The goal of cognitive ergonomics is to improve task performance by the systematic study of the interaction between human cognitive functioning and the systems or environments in which tasks are performed.

A human factors approach to understanding issues of reliability, validity, proficiency, expertise, and sources of bias involves an examination of sources of information that are both internal and external to the forensic examiner. It is therefore important to address the production of forensic science from multiple perspectives. For example, the report of the Expert Working Group on Human Factors in Latent Print Analysis presented a taxonomy of errors and human factors which identified multiple potential sources of error [4].

Recognizing that science is produced by humans who are acting within various systems of professional roles and standards, organizations such as the National Institute of Justice, the Department of Justice, the National Science Foundation, and the National Institute of Standards and Technology (NIST, see Organization for Scientific Area Committees, or OSAC)2 have launched programs and initiatives intended to bring together stakeholders to address these issues using an interdisciplinary approach.

In February 2020, NIST published an extensive report prepared by the Expert Working Group for Human Factors in Handwriting Examination, titled Forensic Handwriting Examination and Human Factors: Improving the Practice Through a Systems Approach [5]. Among other recommendations, the report endorsed ongoing interdisciplinary research encompassing expertise from forensic practice, social and cognitive psychology, vision science, and other areas to establish the basis and extent of forensic document examiner (FDE) expertise, and to develop empirically validated and rigorous document examination protocols, measures, and education and training programs that consistently and comprehensively address the knowledge and skills required to establish expertise in forensic fields.

Dyer, Found, and Rogers utilized eye tracking to study visual attention given to signatures by FDEs and a control group, noting eye movement, response time, and opinions. They found that FDE opinions were significantly more accurate than the control group [6]. However, both FDEs and control participants viewed features of signatures similarly, leading the researchers to suggest that differences in accuracy may be due to different cognitive processes for approaching the task. Specifically, improved accuracy by FDEs involved a capacity to process elemental pieces of information into the context of the total or global information inspected in a signature; whilst well motivated control participants more frequently made decisions only based on elemental visual evidence as they terminated decisions when inspecting one potential error. Thus, FDEs showed initial evidence of processing specific evidence in context, which was also supported by control experiments where global information presented tachistoscopically for brief periods that prevented eye movements did permit correct classification of signatures, but accuracy significantly improved when detailed feature information was enabled by permitting the FDEs to move their eyes to inspect detail [6]. A subsequent forensic eye tracking study also showed a capacity of FDEs to reliably discriminate between forged and disguised or simulated signatures [7].

  1. Using the eye-tracking method employed by Dyer and colleagues [6] in their seminal research on visual attention in document examination, our studies further explored how FDEs extract information from handwritten signatures. Our study specifically investigated the relationship between the arrangement of the questioned and known signature specimens, and the utilization of writing features in comparisons of questioned and known signatures.

  2. The placement of the signatures during the examination created a visual context in which the FDE conducts the comparison, as demonstrated in Figure 1.

Figure 1

This “gaze plot” of a recorded comparison shows an example of a visual context effect. Note that the position of the two known signatures beneath the first name (top and bottom left [AGD1]) shows that this FDE fixated to a greater extent on those two known signatures, while the positioning of the other two known signatures under “LaBarbera” (top and bottom right) attracted a greater number of fixations on those known signatures.

We were specifically interested in understanding how the position of the known signatures in relation to the questioned signatures (context effects) were related to the outcome of FDE decisions about whether the questioned signatures were genuine or non-genuine.

Methods

FDE Sampling, Participant Recruiting, and Data Collection Procedures

Forty-nine professional document examiners working in government laboratories or private practice participated in the eye-tracking study.3 The target number of participants was determined by consideration of several factors, including how many participants would be needed to achieve sufficient power for statistical analyses.4 We recruited participants by attending various professional meetings and presenting information about the research, and we also utilized snowball sampling whereby examiners who participated recommended other members of the field whom they felt would also be willing to participate.

Materials and Equipment

Eye tracking devices track and record eye movements as people interact with the physical environment. Information about eye movements and gaze behavior allows researchers to answer questions about how people deploy their attention. For example, eye tracking can inform researchers about how people interact with web pages, with images, and with visual environments in a variety of situations [8, 9].

Eye trackers consist of three principal components. The illuminators (light source) direct near-infra red light to the pupil and cornea. The light creates highly visible reflections which are captured by the second component, the camera. The final component, the image processing algorithms, use the angle between the reflections of the cornea and pupil to calculate a vector. The vector and other features of the gaze reflections are then used to calculate the gaze direction. In this way the eye tracking unit can record what the participant is looking at, how long the participant’s gaze rests in a particular area, the “scan path,” or the sequence of positions where the participant’s gaze rests (“gaze fixations”), and how long the fixation lasts. The raw fixation data is then filtered to display those areas where the eye position remains within a 50-pixel area for a time of greater than 100 MS. All gaze data were collected using Tobii® model T-60 binocular eye tracking systems with 17” thin film transistor (TFT) screens similar to the technology used in cellular phones, and 1280 x 1024 pixel displays (Tobii® Technology, Stockholm, Sweden), and Tobii Studio© software version 2.3.2.

Signature Stimuli

The signature stimuli were prepared to capture several different signature features that might be encountered as part of the FDE caseload. An objective classification system was utilized by the two FDE subject matter experts on our research team based upon the number of allographs (i.e., a letter of an alphabet in a specific shape, such as capital or lower case, italic, or various handwritten forms of the letter that fall within that letter template) that are present and legible within the signature [10, 11]. This classification scheme identified three types of signature: (1) text-based, in which each allograph of the name is clearly written; (2) mixed, in which two or more allographs can be read; and (3) stylized, in which one or fewer allographs are legible. Our FDE subject matter experts determined signature complexity using a theory developed by Found and Rogers [12, 13]. Complexity theory states that as complexity increases, (as referenced by an increase in the number of features within the writing) the likelihood of someone being able to successfully simulate an image decreases [12, 13, 14]. Mohammed provides a detailed description of this method of determining the complexity of a signature, which is based on the evaluation of the number of turning points, line intersections, and retrace strokes in a signature [15].5

Signature Collection Procedure

Fifty signature writers participated in the signature collection process. All signature writers were unpaid volunteers who were over the age of 18 years. All volunteers were recruited by members of the research team from among friends, neighbors, colleagues, or students. None of the volunteers were police officers or investigators, and none of the volunteers had forensic training.

Writers were seated at a desk or table and provided with pen, paper, and backing sheets. Writers were then given four sheets labeled “Normal Signatures” and instructed to write their signature on the page using their normal signature style, as if they were signing a check or signing their name to routine office or schoolwork. Four signatures were collected per page, for a total of 16 genuine signatures. Writers were given one additional sheet of paper labeled “Model Signatures for Forgers,” on which they used the same procedure to produce three model signatures to be used by other writers to create forged specimens.

Simulated signatures were then created by asking each writer except the writer who produced the original signature to simulate the genuine signature of a previous writer. Writers were not allowed to trace any of the known signatures, and that they were not allowed to practice simulating any of the signatures. The 49 simulators each produced three simulations using the genuine signatures provided as a model/guide.

The FDE subject matter experts then used an online random sequence generator6 to selected 19 simulated and 11 genuine signatures. They classified the signatures as text-based, mixed, or stylized, and calculated the complexity of each writer’s signature based on the classification model equations described above. Of these, they selected six signature specimens for each of the 11 selected signature writers, for a total of 66 questioned/known comparisons (22 genuine signatures, 9 disguised signatures, 22 simulated signatures, and 13 traced signatures). Sixty-two of the questioned signatures were randomly selected. The remaining four signatures were genuine writing specifically selected because of the presence of a feature or characteristic variation produced only in one genuine signature specimen produced by the writer (rare accidental or extraneous features). Figure 2 demonstrates an example of an “accidental” characteristic.

Figure 2

Genuine Terry Lu signature with “eyelet” in the T, which appeared only in this genuine questioned signature. The T in the specimen on the right is a two-stroke letter form in which the staff and the crossbar are not connected.

Procedure

The protocol consisted of 11 trials, each containing six questioned/known signature comparisons. The eye-tracking system was calibrated to each participant using a 9-point calibration grid. Each participant completed a practice trial for the protocol which included an additional calibration check. The material for each signature comparison was presented in the sequence demonstrated in Figure 3.

Figure 3

An example of the stimulus sequence displayed on the eye-tracking system. Each of the four images in the figure were displayed on separate screen. Stimulus 1 is a “fixation cross” where FDEs were instructed to center their gaze so that each comparison began in the same location. After being displayed for 3 seconds, the screen changed to Stimulus 2, the known signatures. FDEs viewed this screen for as long as they chose. When they were ready to complete the comparison, they indicated so verbally to the research assistant, who then changed the screen to Stimulus 3, a second fixation cross. After 3 seconds, the screen automatically changed to Stimulus 4, the questioned/known comparison. FDEs again viewed this screen for as long as they chose. When they completed each comparison they verbally indicated this to the research assistant, who exited the comparison screen. Fixations were recorded from the time the screen changed to the signatures until the research assistant exited that screen.

A fixation cross displayed for three seconds, then automatically changed to the “known signature” screen. The known signature screen displayed four examples of the signature writer’s true, naturally written signature in the lower half of the screen. FDEs studied the four known signatures before conducting their comparison of the questioned and known signatures. No time limit was given for either their examination of the known signatures or the questioned/known comparisons so that accurate responses were enabled [6]. When the FDEs felt that they were ready to view the next stimulus they indicated this verbally to the researcher, who immediately exited to the next screen. Figure 4 illustrates the position of the FDE in relation to the eye-tracking equipment.

Figure 4

The participant is seated 57 cm away from the eye tracking unit. The operator, seated to the participant’s right, controls the presentation of stimuli and records the participant’s verbal responses. The eye-tracking protocols were conducted in low lighting for maximum pupil dilation, which improves the quality of the gaze data. In this image a gaze plot, or map of the fixation locations and the scan path, is displayed. The image displayed on the operator’s screen is a gaze plot created from filtered fixation data.

After each comparison we asked participants for a decision about whether the questioned signature was genuine, disguised, or simulated, then asked for their degree of certainty using the 9-point scale. If the participant responded that the signature was simulated, we asked them to indicate whether they thought that the simulation was a freehand copy or a tracing, and how confident they were about this opinion on a 4-point Likert-type scale ranging from Not at all Confident to Extremely Confident. All responses were manually recorded by the researcher.

Results

These analyses are based on data from 36 signature comparisons of genuine and simulated signatures, conducted by 48 professional FDEs employed in private practice or government laboratories. Results represent 1,742 comparison decision observations. Of these, 723 (42.5%) observations are for genuine signatures and 1,019 (58.5%) are for simulated signatures. Of the 1,742 observations, 1,065 (61.1%) are for text-based signatures and 677 (38.9%) are for stylized signatures. A total of 729 (41.8%) observations are for high complexity signatures, and 1,013 (58.2%) are for low complexity signatures.

Although many metrics are available from our eye-tracking data, we focus here on the total fixation count (FC) metric for the genuine and simulated signature comparisons. A fixation is defined by the speed of the eye movement and the distance between adjacent data points. This metric measures the number of times the participant fixates in an “area of interest” (AOI).

AOIs for this study were created empirically using two visualizations of the cumulative gaze data for all participants. Figure 5 shows examples of a signature comparison stimulus, a gaze plot, a “heat map,” and AOIs created from the information from the two data visualizations. The gaze plot shows the location of all fixations. The heat map shows a cumulative view of where the fixations are concentrated. Darker areas on the heat map show where the highest visual activity occurred.

Figure 5

Image 1 (upper left) is an example of a signature comparison stimulus displayed on the eye-tracking screen. Image 2 (upper right) is a “gaze plot” visualization of the visual behavior of the FDE. Each dot represents a “fixation” recorded by the system when the FDE’s eye movement drops below the threshold described above. The lines represent visual “saccades,” or eye movement between fixations. The fixations are sequentially numbered by the eye-tracker to enable researchers to track the deployment of cognitive attention resources. Image 3 (lower left) is a “heat map” visualization of the areas that received to most attention from the FDE. Image 4 (lower right) demonstrates the areas of interest (AOIs) developed from the heat map.

Each AOI is identified by a different color. An AOI surrounds the questioned signature and separate AOIs surround the known signatures. Smaller AOIs were created within these five larger AOIs where the heat map and gaze plot indicate that fixation activity has been high. One large AOI encompasses all four known signatures. The largest AOI encompasses the entire stimulus. AOIs allowed us to compare the isolated areas separately from the other stimulus areas. The analyses we discuss here are based on the four AOIs isolating the known signatures.

Using these four AOIs, we created two new additional variables by combining the number of fixations into larger units. One variable isolated the fixations according to whether they occurred in the top two or the bottom two known signatures. The other variable isolated the fixations occurring in the two signatures on the left side or the right side of the stimulus. This results in the three AOI configurations demonstrated in Figure 6.

Figure 6

Configuration 1 includes all four known signatures. In configuration 2 the fixation counts and for known signatures 1 and 2 are combined in AOI Top, and 3 and 4 are combined in AOI Bottom. In configuration 3, fixation counts for signatures 1 and 3 are combined in AOI Left, 2 and 4 are combined in AOI Right.

Utilization of Known Signatures

We performed a non-parametric Friedman test of differences among repeated measures7 to investigate whether the differences in the mean number of fixations in the four known signatures were statistically significant. All analyses were conducted at an alpha level of 0.05. Mean FCs for each of the eight AOIs are presented in Figure 7.

Figure 7

The statistically significant difference in the mean number of fixations demonstrated in this figure are consistent with a top-to-bottom and left-to-right viewing pattern. Fixations are highest in the upper left known signature (Known 1) then in the upper right signature (Known 2). Fixation count in Known 3 is significantly lower than Known 2, but significantly higher than in Known 4. Fixation count is significantly higher in the combined areas of Known 1 and 2 (Top), than in Known 3 and 4 (Bottom). Fixation count is also significantly higher in the combined areas of Known 1 and 3 (Left) than in Known 2 and 4 (Bottom). It is important to note that fixation counts are filtered data that identify where the FDE’s gaze has slowed enough to meet the 50-pixel/100 ms threshold requirement. It should not be inferred from these data that FDEs failed to look at other areas where fixations were not indicated by the filter.

We found that mean FCs varied significantly across Knowns 1-4, in a pattern suggesting a clockwise (left-to-right on top and right-to left on the bottom) deployment of attention due to the lower mean FC in Known 3 than in Known 4. A top-to-bottom pattern is clearly displayed by the higher number of fixations in the top Knowns compared to the bottom Knowns. A left-to-right pattern is also indicated in the higher number of fixations in the left Knowns than in the right. Significant results are highlighted in Table 1.8

Table 1

Mean Fixation Count by Area of Interest

95% Confidence Interval
Fixation Count AOIsNMSDtdfPLower CIUpper CI
FC Knowns 1, 2 (top)165024.4722.9526.921649< .0019.5411.04
FC Knowns 3, 4 (bottom)165014.1819.07
FC Knowns 1, 3 (left)165020.121.54.41649< .0010.862.24
FC Knowns 2, 4 (right)165018.5520.25
Fixation Count AOIsNMSDχ2dfpMean Rank
FC Known 1165213.2913.051713.733< .0013.28
FC Known 2165011.1611.232.95
FC Known 316506.810.511.85
FC Known 416507.3911.331.92

We compared the mean FCs for AOI configuration 2 (top Knowns versus bottom Knowns) and 3 (left Knowns versus right Knowns) by conducting paired-groups t-tests to investigate whether significant differences in mean FCs existed within the total observations for these AOIs.

Figure 8 highlights the AOIs where FCs were highest overall. Examination of the mean FCs from highest to lowest in each configuration reveals the clockwise, left-to-right, and proximal pattern of attention deployment.

Figure 8

AOIs in which the highest mean fixation counts are highlighted. Areas of darker shading indicate AOIs with the higher fixation counts. All mean differences are statistically significant, as indicated in Table 1.

Conclusion: Higher mean FCs in AOIs nearest to the known signature, and higher mean FCs in AOIs on the left side compared to the right side and means that decline in clockwise order indicate that use of the known signatures in this sample of observations reflects a left-to-right and proximal-to-distal pattern of attention.

Visual Behavior and Signature Characteristics

The differences observed above may also be influenced by characteristics of the signatures themselves. In this section we report analyses addressing the relationship between the type of signature (text-based or stylized), signature complexity (high or low complexity), and ground truth (whether the signature is genuine or non-genuine) and the gaze behavior of FDEs. We conducted a series of independent samples t-tests to investigate whether the mean FCs in each AOI differed significantly according the level of each signature characteristic (e.g., text-based vs. stylized, high vs. low complexity, genuine vs. non-genuine).

We also conducted a series of binomial logistic regression analyses to investigate which AOIs in combination might be significant predictors of the characteristics of the signatures (the outcome variables signature type, signature complexity, and ground truth).9

Signature Type Analyses

We performed a series of independent group t-tests to investigate whether the mean number of FCs differed according to whether the signature was text-based vs. stylized. The results are presented in Table 2. Statistically significant comparisons and the highest means are highlighted in gray.

Table 2

Mean Fixation Count by Signature Type

Signature Type
95% Confidence Interval
Sig TypeNMSDtdfpLower CIUpper CI
FC Knowns 1, 2 (top)Text100825.6424.332.70115260.0070.8245.195
Stylized64222.6320.49
FC Knowns 3, 4 (bottom)Text100815.0320.982.42916120.0150.4213.957
Stylized64212.8515.53
FC Knowns 1, 3 (left)Text100821.0822.882.41615350.0160.4744.559
Stylized64218.5619.05
FC Knowns 2, 4 (right)Text100819.6022.112.78815970.0050.7954.568
Stylized64216.9116.81
FC Known 1Text101014.1413.873.46715380.0010.9483.419
Stylized64211.9511.51
FC Known 2Text100811.4811.841.45415080.146-0.2781.874
Stylized64210.6810.19
FC Known 3Text10086.9111.360.57416480.566-0.7371.346
Stylized6426.619.02
FC Known 4Text10088.1212.993.6541646< .0010.8732.896
Stylized6426.247.96
Table 3

Regression Coefficients for Predictors of Signature Type

95% Confidence Interval
Analysis 1: Top/BottomBWalddfpOddsLower CIUpper CI
FC Knowns 1, 2 (top)-0.0051.7910.1810.9950.9891.002
FC Knowns 3, 4 (bottom)-0.0020.27510.60.9980.991.006

These results reveal that in all comparisons but Known 2 and Known 3 the mean FC was higher in the text-based than in the stylized condition.

We performed a series of binary logistic regression analyses to investigate whether the observed differences in FC among the known signatures were related to whether the signatures were textbased or stylized. We analyzed each of the three AOI configurations identified above separately to avoid the influence of multicollinearity, which occurs when variables are highly correlated.10

Analysis 1: AOIs Top/Bottom. We combined the predictor variables FC Top (Knowns 1, 2), and FC Bottom (Knowns 3, 4) in a binomial logistic regression model to test how accurately the two factors together predicted category of signature type (the outcome variable).

The overall model was statistically significant (χ2 (2) = 7.23, p = 0.027), indicating that the variables improved the amount of variability explained (Nagelkerke r2 = 0.59%). Results indicated that the overall model fit was poor (-2 Log Likelihood = 2198.29). The model correctly classified 61.1% of cases. Wald statistics indicated that when the variables were combined, neither predictor variable reached statistical significance.

Analysis 2 (Left/Right). We combined the predictor variables FC Left (Knowns 1, 3) and FC Right (Knowns 2, 4) in a binomial logistic regression model to test how accurately factors together predicted category of signature type (the outcome variable). Results are presented in Table 4.

Table 4

Regression Coefficients for Predictors of Signature Type

95% Confidence Interval
Analysis 2: Left/RightBWalddfpOddsLower CIUpper CI
FC Knowns 1, 3 (left)-0.0020.21210.6460.9980.9911.006
FC Knowns 2, 4 (right)-0.0051.74910.1860.9950.9861.003

The overall model was statistically significant (χ2 (2) = 7.36, p = 0.025), indicating that the variables improved the amount of variability explained (Nagelkerke r2 = 0.60%). Results indicated that the overall model fit was poor (-2 Log Likelihood = 2198.16). The model correctly classified 61.1% of cases. Wald statistics indicated that when the variables were combined, neither predictor variable reached statistical significance.

Analysis 3 (All Knowns). We combined the variables FC Known 1, FC Known 2, FC Known 3, and CF Known 4 in a binomial logistic regression model to test how accurately the four factors together were in predicting category of signature type (the outcome variable). Results are presented in Table 5

Table 5

Regression Coefficients for Predictors of Signature Type

95% Confidence Interval
Analysis 3: All KnownsBWalddfpOddsLower CIUpper CI
FC Known 1-0.02813.5641< .0010.9730.9590.987
FC Known 20.0216.77410.0091.0211.0051.038
FC Known 30.0186.25810.0121.0181.0041.033
FC Known 4-0.0227.86510.0050.9780.9630.993

The overall model was statistically significant (χ2 (4) = 29.47, p = 0.001), indicating that the variables improved the amount of variability explained (Nagelkerke r2 = 2.4%). Results indicated that the overall model fit was poor (-2 Log Likelihood = 2176.05). The model correctly classified 61.2% of cases. Wald statistics indicated that when the variables were combined, all four variables were significant predictors of category of signature type. The odds ratios for all predictors were small, indicating low impact on the outcome.

Conclusion: The mean FC was higher in the text-based than in the stylized condition, although the differences in FCs in Known 2 and Known 3 did not reach statistical significance. Based on these analyses, we concluded that although the mean FCs for AOIs top/bottom and left/right were statistically significant, the known signatures were better predictors of signature type when their influence was measured separately, although analysis 3 explained only 2.4% of the variance. Odds ratios in all cases were very near to 1.0, indicating that change in the probability of the signature being text-based vs. stylized was very small.

Signature Complexity Analyses

We performed a series of independent group t-tests to investigate whether the mean number of FCs differed according to whether the signature was high complexity vs. low complexity. The results are presented in Table 6. Statistically significant comparisons and the highest means are highlighted in gray.

Table 6

Mean Fixation Count by Signature Complexity

95% Confidence Interval
CompNMSDtdfpLower CIUpper CI
FC Knowns 1, 2 (top)High69024.2222.69-0.37616480.707-2.6781.817
Low96024.6523.15
FC Knowns 3, 4 (bottom)High69015.7922.672.75011580.0060.7914.732
Low96013.0315.90
FC Knowns 1, 3 (left)High69020.3322.750.36216480.717-1.7162.494
Low96019.9420.56
FC Knowns 2,4 (right)High69019.6822.361.86613090.062-0.1003.984
Low96017.7418.55
FC Known 1High69013.3513.160.16816500.867-1.1681.386
Low96213.2412.97
FC Known 2High69010.8710.95-0.91416480.361- 1.6120.587
Low96011.3811.43
FC Known 3High6906.9712.240.58616480.558-0.7221.336
Low9606.679.07
FC Known 4High6908.8214.484.0121007< .0011.2543.655
Low9606.368.23

These results reveal that the only significant differences in mean FC were found in AOI Bottom (Knowns 3, 4) and AOI Known 4. In both instances the means were higher when the signatures were high complexity.

We performed a series of binary logistic regression analyses to investigate whether the observed differences in FC among the known signatures were related to whether the signatures were high or low complexity.

Analysis 1 (Top/Bottom). We combined the predictor variables FC Top (Knowns 1, 2), and FC Bottom (Knowns 3, 4) in a binomial logistic regression model to test how accurately the two factors together predicted whether signature complexity was high or low (the outcome variable). Results are presented in Table 7.

Table 7

Regression Coefficients for Predictors of Signature Complexity

95% Confidence Interval
Analysis 1: Top/BottomBWalddfpOddsLower CIUpper CI
FC Knowns 1, 2 (top)0.01414.2811< .0011.0141.0071.021
FC Knowns 3, 4 (bottom)-0.02120.2981< .0010.9790.9710.988

The overall model was statistically significant (χ2 (2) = 23.85, p = 0.001), indicating that the variables improved the amount of variability explained (Nagelkerke r2 = 1.9%). Results indicated that the overall model fit was poor (-2 Log Likelihood = 2194.79). The model correctly classified 60.7% of cases. Wald statistics indicated that when the variables were combined, both predictor variables reached statistical significance. The odds ratio for both variables were small, indicating low impact on the outcome.

Analysis 2 (Left/Right). We combined the predictor variables FC Left (Knowns 1, 3) and FC Right (Knowns 2, 4) in a binomial logistic regression model to test how accurately the factors together predicted category of complexity (the outcome variable). Results are presented in Table 8.

Table 8

Regression Coefficients for Predictors of Signature Complexity

95% Confidence Interval
Analysis 2: Left/RightBWalddfpOddsLower CIUpper CI
FC Knowns 1, 3 (left)0.0062.91310.0881.0060.9991.014
FC Knowns 2, 4 (right)-0.016.32510.0120.990.9830.998

The overall model was statistically significant (χ2 (2) = 6.67, p = 0.036), indicating that the variables improved the amount of variability explained (Nagelkerke r2 = 0.54%). Results indicated that the overall model fit was poor (-2 Log Likelihood = 2236.33). The model correctly classified 59.0% of cases. Wald statistics indicated that when the variables were combined, FC Knowns 2, 4 (Right) reached statistical significance. The odds ratio for this variable was small, indicating low impact on the outcome.

Analysis 3 (All Knowns). We combined the variables FC Known 1, FC Known 2, FC Known 3, and CF Known 4 in a binomial logistic regression model to test how accurately the four factors together were in predicting category of complexity (the outcome variable). Results are presented in Table 9.

Table 9

Regression Coefficients for Predictors of Signature Complexity

95% Confidence Interval
Analysis 3: All KnownsBWalddfpOddsLower CIUpper CI
FC Known 1< .001< .00110.9901.0000.9861.014
FC Known 20.03113.8701< .0011.0311.0151.048
FC Known 30.0030.18910.6641.0030.9901.016
FC Known 4-0.04631.8311< .0010.9550.9390.970

The overall model was statistically significant (χ2 (4) = 45.13, p = 0.001), indicating that the variables improved the amount of variability explained (Nagelkerke r2 = 3.6%). Results indicated that the overall model fit was poor (-2 Log Likelihood = 2197.84). The model correctly classified 61.0% of cases. Wald statistics indicated that when the variables were combined, variables FC Known 2 and FC Known 4 were significant predictors of category of complexity. The odds ratios for all three predictors were small, indicating low impact on the outcome.

Conclusion: Only two means (AOI Bottom and AOI Known 4) reached statistical significance. We observed that when combined with other predictor variables, FC Top and Bottom were significant predictors of the category of signature complexity in model 1. FC Right was also a significant predictor in model 2, and Knows 2 and 4, were significant predictors in model 3.

Based on these analyses, we concluded that again, the known signatures were better predictors of signature type when their influence was measured separately, as demonstrated by the amount of variability explained in Analysis 3. Odds ratios in all cases were still very near to 1.0, indicating that changes in the probability of the signature being high vs. low complexity were very small.

Ground Truth Analyses

We performed a series of independent group t-tests to investigate whether the mean number of FCs differed according to whether the signature was genuine vs. non-genuine. The results are presented in Table 10. Statistically significant comparisons and the highest means are highlighted in gray.

Table 10

Mean Fixation Count by Ground Truth

Ground Truth
95% Confidence Interval
TruthNMSDtdfpLower CIUpper CI
FC Knowns 1, 2 (top)Genuine68628.6425.186.1001291< .0014.8449.437
Non96421.5020.73
FC Knowns 3, 4 (bottom)Genuine68617.6720.656.1721316< .0014.0777.876
Non96411.7017.45
FC Knowns 1, 3 (left)Genuine68624.3224.496.5041214< .0015.0389.391
Non96417.1018.53
FC Knowns 2,4 (right)Genuine68622.0021.045.8111396< .0013.9107.896
Non96416.1019.31
FC Known 1Genuine68815.1714.114.8521331< .0011.9184.521
Non96411.9512.06
FC Known 2Genuine68613.4312.556.7401250< .0012.7495.006
Non9649.559.89
FC Known 3Genuine6869.1012.907.1031056< .0012.8605.042
Non9645.158.01
FC Known 4Genuine6868.579.963.5921648< .0010.9203.132
Non9646.5512.15

These results reveal that all differences in mean FC were statistically significantly higher for genuine signatures than for non-genuine signatures.

We performed a series of binary logistic regression analyses to investigate whether the observed differences in FCs among the known signatures were related to whether the signatures were genuine or non-genuine (ground truth).

Analysis 1 (Top/Bottom). We combined the predictor variables FC Top (Knowns 1, 3) FC Bottom (Knowns 3, 4) in a binomial logistic regression model to test how accurately the three factors together predicted category of ground truth (the outcome variable). Results are presented in Table 11.

Table 11

Regression Coefficients for Predictors of Ground Truth

95% Confidence Interval
Analysis 1: Top/BottomBWalddfpOddsLower CIUpper CI
FC Knowns 1, 2 (top)-0.0074.43210.0350.9930.9860.999
FC Knowns 3, 4 (bottom)-0.0115.99810.0140.9890.9810.998

The overall model was statistically significant (χ2 (2) = 45.54, p = 0.001), indicating that the variables improved the amount of variability explained (Nagelkerke r2 = 3.7%). Results indicated that the overall model fit was poor (-2 Log Likelihood = 2118.03). The model correctly classified 60.2% of cases. Wald statistics indicated that when the variables were combined, FC Top and FC Bottom were significant predictors of category of ground truth. The odds ratios for these variables were small, indicating low impact on the outcome.

Analysis 2 (Left/Right). We combined the predictor variables FC Left (Knowns 2, 4, 6) and FC Right Knowns 1, 3, 5) in a binomial logistic regression model to test how accurately the three factors together predicted ground truth (the outcome variable). Results are presented in Table 12.

Table 12

Regression Coefficients for Predictors of Ground Truth

95% Confidence Interval
Analysis 2: Left/RightBWalddfpOddsLower CIUpper CI
FC Knowns 1, 3 (left)-0.01412.2891< .0010.9860.9780.994
FC Knowns 2, 4 (right)-0.0040.82410.3640.9960.9891.004

The overall model was statistically significant (χ2 (2) = 41.17, p = 0.001), indicating that the variables improved the amount of variability explained (Nagelkerke r2 = 3.8%). Results indicated that the overall model fit was poor (-2 Log Likelihood = 2193.15). The model correctly classified 60.4% of cases. Wald statistics indicated that when the variables were combined, only FC 1, 3 (left) was a significant predictor of category of ground truth. The odds ratios for this variable was small, indicating low impact on the outcome.

Analysis 3 (All Knowns). We combined the variables FC Known 1, FC Known 2, FC Known 3, FC Known 4, FC Known 5, and FC Known 6 in a binomial logistic regression model to test how accurately the six factors together were in predicting category of ground truth (the outcome variable). Results are presented in Table 13.

Table 13

Regression Coefficients for Predictors of Ground Truth

95% Confidence Interval
Analysis 3: All KnownsBWalddfpOddsLower CIUpper CI
FC Known 1 (top left)0.0218.34110.0041.0221.0071.036
FC Known 2 (top right)-0.03316.5531< .0010.9670.9520.983
FC Known 3 (bottom left)-0.05030.3591< .0010.9510.9340.968
FC Known 4 (bottom right)0.0123.55010.0601.0121.0001.024

The overall model was statistically significant (χ2 (4) = 84.56, p = 0.001), indicating that the variables improved the amount of variability explained (Nagelkerke r2 = 6.7%). Results indicated that the overall model fit was poor (-2 Log Likelihood = 2155.77). The model correctly classified 62.1% of cases. Wald statistics indicated that when the variables were combined, FC Known 1, FC Known 2, and FC Known 3 were significant predictors of category of ground truth. The odds ratios for these variables were small, indicating low impact on the outcome.

Conclusion: All FC means comparisons reached statistical significance, and in all comparisons were greater for genuine than for nongenuine signatures. We observed that when combined with other predictor variables, FC Top and Bottom were significant predictors of the category of signature complexity in model 1. FC Left was also a significant predictor in model 2, and Knowns 1, 2, and 4 were significant predictors of the category of ground truth in model 3.

Based on these analyses, we again concluded that the known signatures were better predictors of signature type when their influence was measured separately, as demonstrated by the amount of variability explained in Analysis 3 (6.7%). Odds ratios in all cases were still very near to 1.0, indicating that changes in the probability of the signature being genuine vs. non-genuine were very small.

FDE Decision Accuracy Analyses

We performed a series of independent group t-tests to investigate whether the mean number of FCs differed according to whether the FDE decisions were accurate vs. misleading. The results are presented in Table 14. Statistically significant comparisons and the highest means are highlighted in gray.

Table 14

Mean Fixation Count by FDE Decision Accuracy

95% Confidence Interval
AccuracyNMSDtdfpLower CIUpper CI
FC Knowns 1, 2 (top)Accurate141322.9022.26-6.380303< .001-14.315- 7.566
Mislead23733.8424.78
FC Knowns 3, 4 (bottom)Accurate141313.2119.06-5.0791648< .001-9.354-4.142
Mislead23719.9618.15
FC Knowns 1, 3 (left)Accurate141318.7121.07-6.234311< .001-12.699-6.606
Mislead23728.3722.22
FC Knowns 2,4 (right)Accurate141317.4019.96-5.573315< .001-10.874- 5.199
Mislead23725.4320.64
FC Known 1Accurate141412.4412.72-6.135308< .001-7.773-3.998
Mislead23818.3313.85
FC Known 2Accurate141310.4510.87- 5.844300< .001- 6.665-3.307
Mislead23715.4312.36
FC Known 3Accurate14136.2610.44-5.057321< .001-5.137-2.259
Mislead2379.9610.41
FC Known 4Accurate14136.9511.50-4.284352< .001-4.451- 1.650
Mislead23710.009.90

These results reveal that all differences in mean FC were statistically significantly higher for misleading than for accurate FDE decision.

We performed a series of binary logistic regression analyses to investigate whether the observed differences in FCs among the known signatures were related to whether the FDE decisions were accurate or misleading.

Analysis 1 (Top/Bottom)

We combined the predictor variables FC Top (Knowns 1, 2), and FC Bottom (Knowns 3, 4) in a binomial logistic regression model to test how accurately the two factors together predicted category of decision accuracy (the outcome variable). Results are presented in Table 15.

Table 15

Regression Coefficients for Predictors of Decision Accuracy

95% Confidence Interval
Analysis 1: Top/BottomBWalddfpOddsLower CIUpper CI
FC Knowns 1 and 2 (top)0.01718.0151< .0011.0171.0091.024
FC Knowns 3 and 4 (bottom)< .0010.00210.9611.0000.9911.009

The overall model was statistically significant (χ2 (2) = 38.27, p = 0.001), indicating that the variables improved the amount of variability explained (Nagelkerke r2 = 4.1%). Results indicated that the overall model fit was poor (-2 Log Likelihood = 1319.16). The model correctly classified 85.5% of cases. Wald statistics indicated that when the variables were combined, FC Top was a significant predictor of category of decision accuracy. The odds ratios for FC Top was small, indicating low impact on the outcome.

Analysis 2 (Left/Right)

We combined the predictor variables FC Left (Knowns 1, 3) and FC Right (Knowns 2, 4) in a binomial logistic regression model to test how accurately the factors together predicted category of decision accuracy (the outcome variable). Results are presented in Table 16.

Table 16

Regression Coefficients for Predictors of Decision Accuracy

95% Confidence Interval
Analysis 2: Left/RightBWalddfpOddsLower CIUpper CI
FC Knowns 1 and 3 (left)0.0138.10110.0041.0131.0041.022
FC Knowns 2 and 4 (right)0.0051.35210.2451.0050.9961.015

The overall model was statistically significant (χ2 (2) = 35.18, p = 0.001), indicating that the variables improved the amount of variability explained (Nagelkerke r2 = 3.8%). Results indicated that the overall model fit was poor (-2 Log Likelihood = 1322.80). The model correctly classified 85.3% of cases. Wald statistics indicated that when the variables were combined, the only predictor variable to reach significance was FC Left. The odds ratio for FC Left was small, indicating low impact on the outcome.

Analysis 3 (All Knowns)

We combined the variables FC Known 1, FC Known 2, FC Known 3, and FC Known 4 in a binomial logistic regression model to test how accurately the four factors together were in predicting category of decision accuracy (the outcome variable). Results are presented in Table 17.

Table 17

Regression Coefficients for Predictors of Decision Accuracy

95% Confidence Interval
Analysis 3: All KnownsBWalddfpOddsLower CIUpper CI
FC Known 10.0163.77610.0521.0161.0001.033
FC Known 20.0173.06310.0801.0170.9981.036
FC Known 30.0050.37410.5411.0050.991.019
FC Known 4-0.0040.27210.6020.9960.9811.011

The overall model was statistically significant (χ2 (4) = 39.39, p = 0.001), indicating that the variables improved the amount of variability explained (Nagelkerke r2 = 4.2%). Results indicated that the overall model fit was poor (-2 Log Likelihood = 1318.60). The model correctly classified 85.3% of cases. Wald statistics indicated that when the variables were combined, none of the predictor variable reached significance.

Conclusion: Although all mean comparisons were statistically significantly higher for misleading calls, none of the individual known signatures were significant predictors of the category of decision accuracy. Analysis 1 revealed that FC Top significantly predicted decision accuracy, and Analysis 2 revealed the FC Left was also a significant predictor. All three models explained approximately the same amount of variability. Taken together, these analyses suggest that Known 1, which is common to both AOI Top (Knowns 1, 2) and AOI Left (Knowns 1, 3), accounts for these statistically significant differences.

Combined Signature Characteristic and Fixation Count Analyses

We conducted chi-square2) analyses to determine whether significant differences in FDE decision accuracy occurred for text-based versus stylized signatures, high complexity versus low complexity signatures, and genuine versus nongenuine signatures. The difference in decision accuracy for text-based (86.4%) versus stylized signatures (85.4%) was not statistically significant, (χ2 (1) = .349, p = 0.554). Decision accuracy was significantly greater for high complexity (89.8%) than for low complexity signatures (83.2%), (χ2 (1) = 262.92, p = 0.001). Decision accuracy was significantly greater for non-genuine (97.4%) than for genuine signatures (69.9%). This means that in this sample of signatures, both complexity and ground truth are significant predictors of FDE decision accuracy.

As a final step, we combined all the predictor variables (signature type, signature complexity, ground truth, and all FC variables) in a binomial logistic regression model to investigate which factors together predicted the category of decision accuracy (the outcome variable). Because of high multicollinearity (correlations among the FC variables), we again included only the individual signature variables (FC Known 1, FC Known 2, FC Known 3, and FC Known 4) in the final decision accuracy analysis. Results are presented in Table 18.

Table 18

Regression Coefficients for Decision Accuracy Predictors

95% Confidence Interval
Predictor VariablesBWalddfpOdds RatioLower CIUpper CI
Signature Type-0.0901.29310.2560.9140.7821.068
Complexity0.2291.87810.1711.2580.9061.746
Ground Truth-2.699153.0351< .0010.0670.0440.103
FC Known 10.0278.12510.0041.0271.0081.047
FC Known 2< .001< .00110.9891.0000.9801.021
FC Known 3-0.0101.28110.2580.9900.9731.007
FC Known 40.0030.08210.7751.0030.9831.023

The overall model was statistically significant (χ2 (7) = 292.17, p = 0.001), indicating that the variables improved the amount of variability explained (Nagelkerke r2 = 28.9%). Results indicated that the overall model fit was poor (-2 Log Likelihood = 1065.81). The model correctly classified 85.5% of cases. Wald statistics indicated that when the variables were combined, the only significant predictors of category of decision accuracy were ground truth and FC Known 1. Although this model explained considerably more of the variability (28.9%) than any of the previous models, the odds ratios for the significant predictor variables were still small, indicating low impact on the likelihood of predicting the decision accuracy outcome.

Summary and Conclusions

Use of Known Signatures. We investigated FDE use of known signatures by comparing the mean FCs for three different AOI configurations. FCs were higher for known signatures nearest to the questioned signature. Signatures on the left attracted more fixations than did the signatures on the right. Higher mean FCs in AOIs nearest to the known signature, higher mean FCs in AOIs on the left side compared to the right side, and means that decline in clockwise order indicate that use of the known signatures in this sample of observations reflects a left-to-right and proximal-to-distal pattern of attention during the comparisons.

Use of Known Signatures by Signature Type. We investigated the relationship between several signature characteristics and FDE use of the known signatures by comparing mean FCs and conducting binomial logistic regression analyses for each of the three different AOI configurations. These analyses revealed that the mean FC was higher in the text-based than in the stylized condition for all comparisons but FCs in Known 2 and Known 3. Although the mean FCs for AOIs top/bottom and left/right were statistically significant, the known signatures were better predictors of signature type when their influence was measured separately.

Use of Known Signatures by Signature Complexity. We investigated FDE use of known signatures when the signatures were high complexity versus low complexity. We again concluded that the known signatures were better predictors of signature type when their influence was measured separately, as demonstrated by the amount of variability

Use of Known Signatures by Ground Truth. We performed an analysis of means and a series of binary logistic regression analyses to investigate whether the observed differences in FCs among the known signatures were related to whether the signatures were genuine or non-genuine (ground truth). All FC means comparisons reached statistical significance, and in all comparisons the means were greater for genuine than for non-genuine signatures. We again concluded that the known signatures were better predictors of signature type when their influence was measured separately, as demonstrated by the amount of variability explained in the regression model including only the individual known signatures (6.7%).

Use of Known Signatures and FDE Opinion Accuracy. We performed a series of binary logistic regression analyses to investigate whether the observed differences in FCs among the known signatures were related to whether the FDE decisions were accurate or misleading (decision accuracy). Although all mean comparisons were statistically significantly higher for misleading calls, none of the individual known signatures were significant predictors of the category of decision accuracy. All three regression models explained approximately the same amount of variability (4.1%). Taken together, these analyses suggest that Known 1, which is common to both AOI Top (Knowns 1, 2) and AOI Left (Knowns 1, 3), accounts for these statistically significant differences.

Signature Characteristics and FDE Decision Accuracy. We conducted chi-square2) analyses to determine whether significant differences in FDE decision accuracy occurred for text-based versus stylized signatures, high complexity versus low complexity signatures, and genuine versus non-genuine signatures. Decision accuracy was significantly related to signature complexity and ground truth, but not significantly related to signature type.

Decision Accuracy, Signature Characteristics, and Known Signature AOI Fixation Count. After combining all signature characteristic variables and the AOI variables FC Known 1, FC Known 2, FC Known 3, and FC Known 4 in a final binary logistic regression model to investigate which of these variables were significant predictors of the category of FDE decision accuracy. Taken in combination, the only significant predictors of FDE decision accuracy were ground truth and FC Known 1 (upper left).

Implications

The movement of expert testimony from the status of “proffer” to that of “admissible evidence” is a social process in which experts, attorneys, and judges all participate. It is a negotiated movement from “science,” which is itself a social construction [18], to “legal science” [19], which is mediated by the rhetoric and discourse of attorneys, judges, and academicians. Transparency of methods is an important component of the admissibility of FDE testimony. Eye tracking methodology, physiological data, the diagnostic value of the evidential features of handwriting, and descriptions of the decision-making process will help increase the transparency of the examination process, improving the quality of performance of attorneys, judges, and experts.

Practitioners from many forensic fields have taken seriously the need for standardized training and proficiency testing, and through organizations such as the NIST OSAC are working nationally and internationally to define and establish valid and reliable measures of certainty, proficiency, and error. Forensic experts around the world are striving to ensure that their methods are transparent to the courts, and that judges are given the information they need to make their decisions. Efforts to organize and present information effectively, which are important goals of OSAC, have been an important consequence of the Daubert trilogy11 and the 2009 National Academies of Science report [1]. Forensic scientists are also seeking opportunities to collaborate with judges, attorneys, and scientists from other fields on research and education projects.

These findings have implications not only for Questioned Document Examination, but also for other areas of Pattern and Physics Evidence identified by NIST OSAC and other organizations. Our research methods can be adapted to other disciplines, which will increase the understanding of cognitive human factors in those fields and provide information about possible sources of cognitive bias, such as the visual or cognitive context of the examination, top-down/bottom-up processing of information, order of presentation effects, word-superiority effects, and other relevant cognitive phenomena.

Limitations and Future Directions

Although these findings are based on 1,742 FDE decisions that afford good statistical power, these experimental findings do not approximate the working document examination laboratory conditions under which signature comparisons are conducted. The participants were constrained by the experimental environment and the amount of control necessary to ensure that possible confounding circumstances were held as constant as possible. FDEs were limited to only four known signatures for their comparisons, and they did not have access to the tools or resources that are typically available to them during an examination. Future research could address not only the underlying cognitive and physiological aspects of the examination process, but also the broader laboratory environment of the examiner, where the examinations really occur.

Conclusion

Despite these constraints, experimental data provide valuable insights about the mechanisms underlying the decision-making process. Eye movement studies about human gaze behavior have revealed that visual behavior changes depending on the visual task, whether we are casually interacting with our natural environment, whether our attention is captured by salient features, or whether we are engaging in goal-directed behaviour in which our attention is strategically focused on what is relevant to accomplishing the tasks we perform [20, 21, 22].

The results of this study are consistent with the findings of these and other studies, as this study demonstrated that context effects do impact the visual behavior of FDEs. Mitigating any misleading effects due to the relative positions of the questioned and known signatures may be as simple as shifting the position of the questioned and known materials during the examination to ensure that all the available data are utilized. Better understanding the physiological and cognitive contexts of document examination offers important opportunities to find potential sources of bias and error, and to develop evidence-based protocols and procedures to help mitigate their influence. Ongoing research in both experimental academic laboratories and forensic laboratories where examinations occur is needed to gain a full understanding of the relationship between human factors and forensic science.

Notes

[1] This research was supported by NIJ Award No. 2010-DN-BX-K271, awarded by the National Institute of Justice, Office of Justice Programs, U.S. Department of Justice. The opinions, findings, and conclusions or recommendations expressed in this publication/program/exhibition are those of the author(s) and do not necessarily reflect those of the Department of Justice.

[4] The data for one participant had to be excluded from the study due to a research protocol violation.

[5] Power analysis using techniques described by Cohen and Cohen (1983) were used to determine the number of observations needed per variable (120) to detect a medium effect size of 0.60 at a power of 0.80. The number of observations in each comparison reported here far exceeds the number required to reach this level of power.

[6] Turning points are defined as the number of direction changes and the number of starting points and terminating points of any continuous line [16, 17]. Line intersections or retraced strokes are determined by counting the number of times a line either intersects or retraces a previously written stroke.

[8] The Friedman test is a non-parametric equivalent of a one-sample related measures design. Test variables are ranked (in this case from 1 to 4). The chi-square test statistic is based on the ranks.

[9] N = the number of observations on which the calculations are based. N for these observations is lower than N for the accuracy analyses because eye tracking data are not available for every examiner. We were unable to record gaze data for three FDEs whose corrective lenses interfered with the infrared illumination. These FDEs completed the comparisons and gave opinions, but no gaze data were collected. In several instances the eye tracking unit lost connection during a comparison, resulting in loss of an observation for that comparison. M = mean, and SD refers to the standard deviation of observations from the mean.

[10] While the t-tests measure whether the two group means are significantly different, a binomial logistic regression combines “predictor” variables (in this case, AOIs) in a model to predict which category of the two-category outcome variable (for example, high or low complexity) the observations fall under. The higher the number of correctly classified observations, the more reliable the model is considered. While the t-tests can tell us which AOIs are significantly different, they do not tell us anything about how the AOIs work in combination during an examination. Binomial logistic regression analyses reveal the changes in the odds of the observation falling into one category or the other, allowing us to identify the most influential predictor AOIs.

[11] For example, AOI Top is a combination of AOI Knowns 1 and 3, so is highly correlated with Known 1 and Known 3.

[12] Daubert v. Merrell-Dow Pharmaceuticals, Inc., 509 U.S. 579 (1993). The four pronged “test” for the evidence/expert includes the following guidelines: (1) general acceptance, (2) falsifiability, (3) error rate, and (4) peer review and publication. Daubert also gave birth to two other controversial and historical cases that with Daubert, became known as the “Daubert Trilogy”: Gen. Elec. Co. v. Joiner, 522 U.S. 136 (1997) and Kumho Tire Co. v. Patrick Carmichael, 526 U.S. 137 (1999).

Acknowledgements

NIJ Award No. 2010-DN-BX-K271: Kentucky State University: Undergraduate assistants Sandra Keene, Mickail Buster-Jones, Inna Malyuk, Cierra Alexander, Kara Francis, Nick Williams, Laurice Jackson, La’Quida Smith, Ivan Duvall, Savada Smothers, and Melissa Pickett; University of Nevada, Reno: Christopher Sanchez, Survey Manager; Martha Rodriguez, Assistant Survey Manager; Christopher Swinger, Supervisor; and the survey interview staff of the Center for Research Design and Analysis; Denise Schaar-Buis, Research Faculty Assistant, Grant Sawyer Center for Justice Studies; and graduate assistants Mauricio Alvarez, Lindsay M. Perez, and Vicky Springer of the Grant Sawyer Center for Justice Studies, University of Nevada, Reno.

DOI: https://doi.org/10.69525/jasqde.262 | Journal eISSN: 1524-7287
Language: English
Page range: 9 - 30
Published on: Jun 1, 2020
Published by: American Society of Questioned Document Examiners
In partnership with: Paradigm Publishing Services

© 2020 Mara L. Merlino, Veronica B. Dahir, Charles P. Edwards, Derek L. Hammond, Tierra Freeman-Taylor, Adrian Dyer, Bryan J. Found, published by American Society of Questioned Document Examiners
This work is licensed under the Creative Commons Attribution 4.0 License.