Deep learning artificial intelligence models compared to ordinary linear regression for prediction of visual field progression
Highlight box
Key findings
• All neural network models we tested were found superior to ordinary least squared regression when predicting visual field progression, but no one model was superior to the rest.
What is known and what is new?
• It is known that prediction of visual field loss with the present standard analysis of ordinary least squares method has significant limitations as it only uses the mean deviation summary measure
• This study compares the spatio-temporal prediction from a number of different well described machine learning (ML) models which learn from all the sensitivity data to predict progression of visual field loss.
What is the implication, and what should change now?
• Our study proves the utility of using the complete sensitivity data with ML rather than regressing a summary measure, but an optimal model has yet to be found.
Introduction
Assessment of functional damage to the field of vision (VF) is essential for the diagnosis, treatment, and management of many diseases of the eye and visual pathway (1). A number of the retinal and nerve related diseases affecting visual field integrity may progress over time, with most having a characteristic pattern of field loss. Typically, optic nerve or cerebral vascular occlusions (‘strokes’) produce acute characteristic changes to the visual field which remain static. Glaucoma, on the other hand, is usually a slowly developing idiopathic disease affecting the optic nerve head which progresses to a profound irreversible blindness if left untreated. With glaucoma the retinal nerve fibres are damaged resulting in the loss of ganglion cells. The field defects produced usually start as subtle focal losses of light sensitivity that develop into characteristic patterns of field loss overtime.
Standard automated perimetry is an integral part of ophthalmic assessment and monitoring of visual function. The VF test with the Humphrey Field Analyser (HFA) using the 24-2 test pattern produces a 54-location representation of a single eye’s central 24-degree visual field, often referred to as a monocular visual field (1). Measures of the threshold light difference between the static target and background, in decibels, represent the retinal sensitivity at each location. It provides a spatial representation of an eye’s VF, giving a range of values from no vision [0 decibel (dB), no response to the maximal 10,000 apostilb (Asb) light stimulus] to very good vision (40 dB, a response to a 1 Asb stimulus target to background differential) at each of the 54 locations. As the test is subjective, that is, requiring active patient participation, reliability depends upon patient factors such as alertness and mood. The monocular 24-2 Humphrey visual field (VFH) is the predominant visual field test performed in clinics for the detection and diagnosis of various problems, in particular glaucoma (1).
Research around visual defect progression is dominated by progressive glaucomatous disease, with ophthalmologists using a mixture of subjective evaluations of the visual field sequence and the software provided by the manufacturer of the machine (2). The European Glaucoma Society practice guidelines recommend three visual fields in the first year to obtain a baseline once the patient is diagnosed (3). There is a learning curve for the patient to understand how to respond correctly to the stimuli in a standard automated perimetry test, so around 10% of patients, may show a slight decrease in their visual field sensitivity in their first test results (4-7).
Following the Early Manifest Glaucoma Trial (8) criteria, the Zeiss Humphrey Glaucoma Progression Analysis Change Probability Maps use two baseline tests, then three follow-up tests to determine ‘progression’ in a patient’s disease. Field defect progression occurs if there is a pattern deviation in the same three or more locations in three consecutive visual field tests (9). That requires a mean time of 33 months to detect confirmed progression (10).
Currently, it takes time to determine if a patient’s visual field defect has progressed, so local methods or global indices are generally used. Both of those, however, do not address any spatial relationships and are unable predict change over time (1).
Some machine learning (ML) models to detect glaucomatous progression have been assessed to date. Compared with the pointwise linear regression (PLR) method that is included the Humphrey Machine statistical analysis package (STATPAC2) a variational Bayesian Independent Component Analysis mixture model (vB-ICA-mm) (11,12) and variational Bayesian Mix Factor Analysis (vB-MFA) (13) produced good clustering. Whilst such statistical analysis at each locus provides temporal information, i.e., change over time, the specific ‘pattern’ of the defect, i.e., diagnosis or type of change, is not made evident.
There are few deep learning (DL) neural network methods applied to visual field prediction published to date (Table 1) (14-18). As the problem of predicting visual field has both a spatial and a temporal nature, not all DL methods are able to solve both. Recurrent neural networks (RNN) are known for solving temporal problems (15), whereas convolutional neural networks (CNNs) are known for solving spatial problems (19). The variational auto-encoder (VAE), on the other hand, is a generative model that learns the complex distributions of high-dimensional data by projecting it into a lower-dimensional latent space, then using that to map it back to the original dimension (14).
Table 1
| Study | Model (spatial-temporal) | Input (blind spot removal) | Constraint | Sequence time | Prediction field | Outcome measure: DL model | Outcome measure: linear regression |
|---|---|---|---|---|---|---|---|
| Present study: 1,526 patients with 6 or more tests from 9,569 patients | Various | 54—all fields (no) | None | Any | 4th, 5th, 6th | RMSE: 2.38 (±1.56 dB) | RMSE: 4.476 (±2.82 dB) |
| Berchuck et al. 2019 (14): all 3,832 patients | VAE (yes) | 52—glaucoma patients only (yes) | Fixation loss <33%; false positives <15%; data normalised | Any | 4th, 6th, 8th | MAE: 5.14 dB | MAE: 8.07 dB |
| Park et al. 2019 (15): 841 patients with 6 or more | LSTM (no) | 52—glaucoma, reliability constraint (yes) | Fixation loss, false positives, and false negative <33% data normalised | Any | 6th | RMSE: 4.31 (±2.54 dB) | RMSE: 4.96 (±2.76 dB) |
| Wen et al. 2019 (16): all 4,875 patients | CNN (no) | 54—all fields (no) | No constraints | >9 months | 4th | RMSE: 3.47 dB (95% CI: 3.45–3.49); MAE: 2.47 dB (95% CI: 2.45–2.48) | MAE: 3.29 (95% CI 3.24–3.34 dB) |
Data are presented as mean (± SD) unless otherwise indicated. CI, confidence interval; CNN, convolutional neural network; DL, deep learning; LSTM, long short-term memory; MAE, mean absolute error; RMSE, root mean square error; SD, standard deviation; VAE, variational auto-encoder.
The main goal of our study was to examine if several commonly used DL models may better predict a sequence of standard automated perimetry results using a real-world clinical dataset from a range of patients. We present this article in accordance with the TRIPOD+AI reporting checklist (available at https://aes.amegroups.com/article/view/10.21037/aes-25-38/rc).
Methods
We evaluated nine different DL neural networks to compare with ordinary least squares regression analysis prediction of the 4th, 5th and 6th visual field dB score at each test loci from a series of 17,880 sequential visual fields ranging from 6 to 14 tests, averaging 261 days between tests (Figure 1). The prediction was based on the first three tests in the series.
To evaluate temporal-spatial changes, we compared RNNs, convolutional networks, and hybrid convolutional recurrent networks. The recurrent networks are of three types, including a simple Elman RNN, Long Short Term Memory (LSTM) recurrent network, Gated Recurrent Units (GRU), also bidirectional versions of RNN, LSTM and GRU (20-22). We used a modified 3-dimensional CNN (C3D) and looked at a hybrid convolutional LSTM (ConvLSTM) with either a single layer or double layer (ConvLSTM) (23,24).
Data extraction, preparation and transformation
The data set from the specialist ophthalmic clinic had patient results from a variety of tests results such as 10-2, 24-2 and 30-2 threshold test. A total of 43,230 visual fields were obtained from 9,569 patients from two HFA machines using the standard recommended clinical test protocol (1). In addition, a total of 38,086 visual fields from 7,872 patients originated from an SN750 machine and 5,144 visual fields from 1,697 patients were from an SN745 machine.
Only the results from 24-2 were utilised in this study; they were the majority of the tests in the whole data set extracted. The raw decibel values for the 54 field loci were extracted for each test producing a string of variables with patient, eye, and date variables appended.
To ensure test data quality, we decided to remove the first test performed by the patient which was in line with Fujino et al. (25). The patient’s second test then became their first test in the stack, which will from now on will be referred to as the first test.
The blind spot in the visual field is either at loci 36 or 46. It seems the most common blind spot location is 46. Wang et al. (26) discovered that if the HFA was unable to locate the patient’s blind spot, the HFA returned a default location for where the blind spot is (15°, −1°) for the right eye and (−15°, 1°) for the left. We kept both loci 36 and 46 maintaining 54 loci for each test.
We then selected patients who had between 6 and 14 tests. Any patients that fell outside of this range were not included. The results were then filtered into six sequential test windows. For example, if a patient had 10 tests: 1–6, 2–7, 3–8, 4–9 and 5–10 were all used as individual sequences. If a patient had 9 tests, then only the following were considered: 1–6, 2–7, 3–8, 4–9 (Table 2).
Table 2
| Test range | Right eye | Left eye |
|---|---|---|
| 1–6 | 432 | 443 |
| 2–7 | 327 | 336 |
| 3–8 | 237 | 247 |
| 4–9 | 162 | 174 |
| 5–10 | 107 | 121 |
| 6–11 | 75 | 81 |
| 7–12 | 49 | 55 |
| 8–13 | 37 | 42 |
| 9–14 | 28 | 27 |
| Total | 1,454 | 1,526 |
The left eye data was mirrored so all the data was configured as a right eye for ease of analysis. The final selected data set totalled 17,880 perimetry results from 2,980 sequences. This was split 80% into the training set and 20% into the test set producing a training set had 14,304 visual fields test results from 2,384 sequences, and a test set of 3,576 visual fields from 596 sequences (Figure 1, Table 3). The characteristics of the training and test set are shown in Figure 2 and Table 4.
Table 3
| Sequential test results | Number of perimetry tests | Number of sequences |
|---|---|---|
| Total | 17,880 | 2,980 |
| Training set | 14,304 | 2,384 |
| Test set | 3,576 | 596 |
Table 4
| Data set | Age, years | Total visits | Mean deviation in dB | 2nd visit mean deviation in dB | AGIS score† |
|---|---|---|---|---|---|
| Training set | |||||
| Mean (SD) | 68.54 (10.42) | 12.86 (4.45) | −6.10 (5.83) | −6.23 (6.11) | 4.17 (4.62) |
| Min – Max | 23.7–89.8 | 8.00–33.00 | −29.62 to 17.81 | −27.42 to 3.67 | 0–20 |
| Test set | |||||
| Mean (SD) | 68.54 (10.34) | 12.61 (4.31) | −5.81 (5.79) | −6.23 (6.11) | 3.88 (4.56) |
| Min – Max | 35.5–89.8 | 8.00–33.00 | −27.37 to 17.81 | −27.42 to 9.85 | 0–20 |
†, AGIS visual field defect score is based on the number and depth of clusters of adjacent depressed test sites in the upper and lower hemifields. AGIS, Advanced Glaucoma Interventional Study; Max, maximum; Min, minimum; SD, standard deviation.
For the convolutional models, the data was reconstructed as an array within a 12×12 matrix with padding using zeros in loci outside the 54 visual field loci. For the other models the data remained as a string (i.e., 144×1 matrix).
Outcome measures
Two metrics are widely used to understand model evaluation: the root mean square error (RMSE) and mean absolute error (MAE). These metrics were used on the test set to understand the difference in the errors produced by the models. MAE assesses the absolute difference between the predicted and actual observation, in contrast the RMSE assesses the square root of the average of the square differences of the predicted and actual observation (27). MAE looks at the average absolute error in the set of predictions, but it does not take into consideration whether the error is positive or negative. Each of the individual errors has equal weight. This is different to RMSE where the square root of the average squared difference between actual and observation is used. RMSE is sensitive to outliers and penalises large errors, whereas MAE weighs the errors similarly, so this means it can reject some large errors.
The visual field prediction that provides the shape of the evolving visual defect is important when assessing the two metrics. Defects can first appear focally and subtly (16) and progress into known shapes over time (28,29). It is because of this, a model will need to be penalised for large errors when producing a prediction, rather than how many overall errors there were. Large errors in a location point prediction may affect the shape of the visual field defect.
Statistical analysis
The null hypothesis for Wilcoxon signed-rank test is whether the median difference between the observation pair is zero. We used this to compare the errors from linear regression and the DL methods. Results obtained with linear regression were considered statistically different at P<0.01.
Models
Traditional data mining techniques cannot perform well on spatio-temporal data for three reasons. Firstly, the data is usually embedded in continuous space whereas the data used in traditional data mining techniques are discrete. Secondly, the data have both spatial and temporal aspects to it, making it more complex. Lastly, correlations amongst the spatio-temporal data is more difficult to capture. A spatial-temporal observation is when a spatial and temporal phenomenon exists at a certain time and location. If a dataset is collected across time as well as space, and it has at least one spatial and one temporal property, such data can then be considered for spatio-temporal modelling. The modelling of visual field prediction falls into this category because when a patient records are collated over time to help medical diagnosis. In addition to this, the visual field data can be considered spatial due to shape and the positional information of each measurement point in the eye. Prior research has focused on either aspect rather than both, with the exception of Berchuck et al. (30).
RNN models
Recurrent units have demonstrated success in labelling and predicting sequential data, especially in natural language processing and speech recognition. The properties of this model being robustness to warping and flexible use of context make them applicable to multidimensional data such as images and video processing. They were originally designed for one-dimensional sequences, however when multiple perimetry tests are carried out in a certain time frame, these results can be treated as a sequence of two-dimensional (2D) arrays, thus making sequential modelling techniques an option. Each type of the recurrent units (RNN, LSTM and GRU) were utilised in predicting visual fields. We started with a single layer of 100 of the respective cells, this was then connected to a layer of 144 dense neurons to then predict each visual field location. An illustration of this is in Figure 3.
One of the problems with recurrent networks is that they do not take into consideration future context. To form a bidirectional RNN, each training sequence forward and backwards form two separate recurrent hidden layers. Both layers are connected to the same output layer. Bidirectional recurrent unit models outperform unidirectional models when using real-world sequence labelling tasks. With visual field results prior research using recurrent units to predict future had only used unidirectional recurrent units because of concerns that bidirectional models may violate causality which is an issue for financial prediction or robot navigation. When there is no distinction between the past and future inputs that can be an issue, however for problems like handwriting recognition where the data can be divided up into sentences or lines, that is not an issue. We used three perimetry test results which provided a baseline for each patient, which could be seen as the division between past and future inputs. Our bidirectional implementations are illustrated in Figure 4.
CNN models
The architecture of convolution networks (ConvNets) is made up of three main layers: Convolutional Layer, Pooling Layer and a Fully-Connected Layer. CNNs generally consist of alternating convolutional and pooling layers.
The convolutional operation is a linear application of a smaller filter to the larger input. This produces a feature map. The filter shifts across the image at a set stride rate. The pooling layers then come in to downscale the feature maps to retain the salient features. When a convolutional filter is 1×1×1 it will have a single weight for each channel in the input. This filter then acts like a single neuron with an input from the same position in each feature map in the input. When applied with a stride of one, it will produce a feature map with the same width and height as the original image. This essentially turns off the convolutional operation and spatial information is not used, rather it is a linear weighting of the input.
In addition to this, the pooling layers will replace the output of a net at a location, by taking the maximum or an average value of the rectangle neighbourhood. This process provides an approximately invariant representation to small translations of the input which allows the preservation of the location of the feature, rather than knowing the exact pixel location of the feature.
Tran et al. (31) proposed the use of 3-dimensional ConvNets (3D ConvNets) to learn spatio-temporal features. They found that 3D ConvNets were more suitable for spatio-tempral feature learning because 2-dimensional convolution networks (2D ConvNets) are only spatially aware and lose temporal information after every convolution operation. Compared to a 2D ConvNet, they demonstrated that 3D convolution and 3D pooling operations of the 3D ConvNets were superior. Prior work has focused on the use of 2D ConvNets by Berchuck et al. (30) and Wen et al. (16), but at time of writing, no papers had investigated the use of 3D convolutional networks on visual field analysis. Our adapted C3D model is depicted in Figure 5.
To understand how important the ability to decipher spatial information was to predicting visual field in the convolutional models, we ‘turned off’ these features in the C3D CNN and ConvLSTM models. The architecture remained the same in C3D CNN models, except the kernel size was reduced to one, and the pooling size were reduced to 1×1×1. For the ConvLSTM models, we adjusted the kernel size because they do not have pooling layers.
For the CNN models we evaluated the effect of using a wide range of filters [4, 8, 16, 32, 64, 124, 256, 512] keeping the other hyperparameters constant. However, we found that the increase of filter size did not provide extra performance to the model, rather the smaller number of filters provided lower error rate on the validation set.
When the lower filters (16 and below) were used, the RMSE of the validation set was below 2 dB for the first and second visual field prediction. When the filter was above 128, the RMSE increased above 2.605. The change in filter size did not have as larger impact on the third visual field prediction, with the filters below 64 resulting in RMSE between 2.470 and 2.6132. The higher filters (128 and above) RMSE were between 2.982 and 3.024. A potential limitation of our experiment is that the kernel size was not optimised. We chose a kernel size of 3×3 based on the research by Tran et al. (31) which indicated that 3×3 kernel was the most effective for spatial-temporal data. So, we used the smaller filters of 8 and 16 for the ConvLSTM.
Experiment details
Training and test data split
A model’s goodness of fit can be examined by providing it an unseen test sample from the data. This is commonly known as a ‘test set’. However, when modelling time series data, to ensure a temporally valid evaluation, we would like to see how the model performs on an ‘out of sequence sample’ dataset; data that is outside of the temporal sequence that was used for training.
The training data in the present study contained a patient’s visual field test sequence if the total number of visual fields in the sequence were between 6 and 14. However, in the original dataset, some patients had 30 or more visual field tests results over the data collection period. Using those ‘out of sequence’ data allowed us to understand how well each model may generalise. Gathering patients who had 15 to 20 visual field test sequences from the dataset (17 patients, 102 visual field tests) we created an ‘out of sequence’ sample, processing them as we had done for the other data sets. The models were then tested, using the 15th, 16th and 17th visual field test results as the input (the out of sequence data set), for predicting the 18th, 19th and 20th visual field test results.
Hyperparameter grid search
A ‘hyperparameter’ is not a model parameter that needs to be learnt through model training. Rather it is one of the learning algorithm parameters and is set before the training process begins (32). The small number of hyperparameters we assessed were the number of cells (neurons) per layer; the number of epochs; the percentage of dropout; and the filter size for ConvLSTM model. We ran all the hyperparameter evaluations on a 5% subset from validation set, which provided an unbiased evaluation of the fit of the model whilst fine-tuning those hyperparameters. The test set was then utilised to assess the performance of the final models.
Geron (32) noted that finding the right number of neurons per layer is somewhat of a ‘dark art’. He noted that the number of neurons per layer is just one hyperparameter to tune, as is the number of layers in a model. A grid search can be utilised to fine-tune a neural network hyperparameter. We utilised Autonomio Talos (33) to reduce the computation expense as it automates the grid search for Keras models.
The hyperparameters that went through the grid search were the number of cells per layer; the number of epochs [10, 25, 50, 75, 100, 125, 150, 175 and 200]; and the range of cells per recurrent model [10, 50, 75, 100]. The other hyperparameters remained constant (activation of Relu, loss of mean square error and the optimizer as Adam). These options produced 36 different models for investigation. Geron (32) suggested that when the hyperparameter search space is large, it is preferable to use a randomised search instead which is offered by Talos (33). We randomly ran 80% of the 36 potential models comparing the loss of each model. As Talos is unable to provide a plot of the training and validation loss to understand if the model is over-fitting or not, we took selected hyperparameters from the grid search and re-ran them to obtain the training and validation loss plot to double check.
After the loss plots were performed, we chose the final for the single layer recurrent hyperparameters: RNN—100 cells and 200 epochs; GRU—100 cells and 200 epochs; LSTM—100 cells and 200 epochs.
Leveraging off the grid search for the recurrent units, an additional model was created that included a bidirectional layer for each recurrent unit. Additional testing on the effect of dropout was run on each of these models with the validation test set.
Dropout
‘Dropout’ (when the training randomly drops units) was introduced by Srivastava et al. (34) to help prevent over-fitting in neural networks and as a way of regularizing a neural network. In our evaluation, the validation only contained 894 visual field results, so dropout did not provide any improvement but rather, slowly deteriorated the model’s ability to predict. As a result, we did not utilise dropout in any of our models.
Normalisation
‘Normalising’ data (scaling it between zero and one) may help the models converge faster, and it may also assist when data is on different scales (32). A scale of a variable that differs (especially when the scale is greater) from the others can have a larger effect on the model (35). We evaluated each neural network with, and without normalised data using the formula: Xnorm= (X − Xmin)/(Xmax − Xmin) to normalise. There was minimal difference on the RMSE between the two groups, and performance improvements were negligible. The visual field test results are all within the same range (0 to 40 dB), so normalising the data did not seem to assist as the data was not on different scales.
Compute infrastructure
The models were trained using Google Colab, running on Tensorflow with Keras and used Python 3.6. After the models were trained, the data was analysed and visualised on a MacBook Pro with a processor of 2.6 GHz 6 core Intel Core i7 with 16 GB 2,400 MHz DDR4.
Ethical considerations
This study was conducted in accordance with the Declaration of Helsinki and its subsequent amendments. The study was approved by institutional ethics committee of University of Western Australia (Approval #2020/ET000250) and individual consent for this retrospective analysis was waived.
Results
We found all the DL models had a significantly better RMSE than the ordinary linear regression (OLR) predictions (Wilcoxon signed-rank test, P<0.01, Figure 6, Table 5). For the 4th sequential prediction, the modified C3D CNN model was best (2.376 vs. 3.936 dB for the linear regression), for the subsequent 5th and 6th sequential prediction the combined ConvLSTM single layer model was best, although not as good as the prediction for the 4th sequence (2.681 vs. 4.155 dB for OLR and 2.838 vs. 4.476 dB for OLR, respectively). Generally, the prediction degraded the further out the sequence for all models and for the linear regression.
Table 5
| Prediction RMSE results | 4th VFH | 5th VFH | 6th VFH |
|---|---|---|---|
| Simple RNN | 2.841 (1.556) | 2.960 (1.674) | 3.126 (1.986) |
| Simple LSTM | 2.849 (1.375) | 2.944 (1.530) | 3.079 (1.835) |
| GRU | 2.766 (1.450) | 2.873 (1.583) | 3.025 (1.899) |
| Bidirectional RNN | 2.633 (1.342) | 2.779 (1.461) | 2.954 (1.776) |
| Bidirectional LSTM | 2.576 (1.283) | 2.706 (1.421) | 2.891 (1.782) |
| Bidirectional GRU | 2.533 (1.321) | 2.682 (1.445) | 2.866 (1.792) |
| ConvLSTM single layer | 2.579 (1.289) | 2.681 (1.417)† | 2.838 (1.736)† |
| ConvLSTM double layer | 2.581 (1.313) | 2.702 (1.456) | 2.879 (1.784) |
| Modified C3D CNN | 2.376 (1.563)† | 2.961 (1.476) | 3.108 (1.750) |
| Ordinary linear regression* | 3.936 (1.942) | 4.155 (2.212) | 4.476 (2.820) |
Data are presented as mean (± SD). *, all the DL models had a significantly better than OLR with Wilcoxon signed-rank test P<0.01. † indicates the lowest RMSE in the column. From a cohort of 9,569 general clinic patients having VFH examinations, 1,526 patients that had six to fourteen consecutive VFH tests were selected. Sequences of six consecutive tests from one patient over time were configured resulting in final selected data set totalled 17,880 perimetry results from 2,980 sequences. That set was split into 80% for training the models, the remaining 20% ‘test set’ was used for the evaluation. The first three tests in the set were used to provide the prediction for the next three. CNN, convolutional neural network; Conv, convolutional; DL, deep learning; GRU, gated recurrent unit; LSTM, long short-term memory; OLR, ordinary linear regression; RMSE, root mean square error; RNN, recurrent neural network; SD, standard deviation; VFH, Humphrey visual field.
We could not really demonstrate a clear benefit of spatio-temporal models over temporal models in our results, although the C3D CNN (a 3D pooling model) and the ConvLSTM (2D pooling) models had the smallest RMSE. The inclusion of spatial information is important for the 3D model, but much less so for the 2D ConvLSTM as shown by the effect of ‘turning off’ the spatial features in Table 6.
Table 6
| Model | With convolution | Without convolution | ||||
|---|---|---|---|---|---|---|
| 4th VFH | 5th VFH | 6th VFH | 4th VFH | 5th VFH | 6th VFH | |
| ConvLSTM single layer RMSE, dB | 2.579 | 2.681 | 2.838 | 2.570 | 2.674 | 2.835 |
| ConvLSTM double layer RMSE, dB | 2.581 | 2.702 | 2.879 | 2.551 | 2.658 | 2.821 |
| C3D CNN modified RMSE, dB | 2.376 | 2.961 | 3.108 | 14.870 | 14.739 | 14.620 |
C3D, three-dimensional convolution; CNN, convolutional neural network; Conv, convolutional; LSTM, long short-term memory; RMSE, root mean square error; VFH, Humphrey visual field.
Discussion
Our study confirms the work of others showing that DL models outperform OLR in visual field prediction.
Whether the addition of spatial information to a temporal sequence model improves prediction, however, remains unclear. Certainly, we found the standard and commonly used computer-vision DL models to date did not greatly enhance predictability. Other spatial-temporal models of a similar nature that we did not evaluate may provide improved prediction.
Comparison of results with the three other reports using DL models for visual field prediction analysis (14-16) is complicated by the differences in approach (Table 1) and by differences in the outcome measure reported. Also, our study had a wide range of patients, not just those with known glaucoma. There were differences in the linear regression results, but they were generally comparable to our findings, however all our DL models appeared better overall.
Of the temporal sequence models used for visual field loss progression prediction, Park et al. (15) used a recurrent unit network. They tested various combinations of fully connected layers and activation functions concluding that the best combination was a single layer of six LSTM cells using a final dense layer that had 52 neurons. The model inputs were the 52 visual field points as total deviation and as pattern deviation values, along with the false negative, false positive and fixation loss numbers, resulting in 107 ‘features’ for each visual field location. Their model provided each field location a prediction value for the sixth visual field after five consecutive visual field tests. They found that the prediction error was lower than the OLR.
Another temporal model, the deep neural net, Cascade-Net, was proposed by Marquez et al. (19) as an efficient training approach using a layered structure in a bottom-up manner. The algorithm splits the network into layers, training each of the input architecture layers one by one. The training procedure could be considered as several interconnecting single-layer 2D ConvNets (sub-networks) that are trained from the bottom to the top. Wen et al. (16) used this method with unfiltered real-world mixed patient Humphrey fields, finding five layers most suitable. They used transfer learning, taking the training weights from the first temporal interval and applied them to the next, then consecutively applying the weights from the proceeding interval along each of the time periods.
Eslami et al. (17) evaluated the models developed by Park (15) and Wen (16) on their own dataset. Their findings supported the low overall pointwise MAE reported in the original studies, but also showed that both models (15,16) significantly underestimated the progression of VF loss (17). To address the scalability of these classical methods, Berchuck et al. (30) proposed a variational autoencoder (VAE) as an unsupervised generative model to predict a patient’s future visual field (14,30). The variational Bayesian inference basis in the VAE is an extension of the traditional autoencoder that enables sampling from the latent space.
Berchuck et al. (30) had to make two assumptions regarding the data for them to use a VAE. Firstly, observations from patient to patient are independent. Secondly, that the visual field images had a low-dimension representation that could be used in place of the original visual field. The second assumption can be backed by Elze et al. (29) who used statistical dimension reduction techniques of Archetypal Analysis to derive 17 visual field patterns, matching those described by Keltner et al. (28).
In their VAE, Berchuck et al. (14) utilised 2D ConvNets as a two-stage approach. In the first stage, latent features are obtained for use in the second stage, using a longitudinal method such as an auto-regressive model. The advantage to this divided approach is that the latent features can be modelled, and their predictions can be used generate forecasts of the visual fields. It is also useful to use the latent space for clustering the disease status. Their model was able to model the data jointly and learn a latent representation of the visual fields that can be interpreted and allowed for a prediction. Using three visual fields, they predicted the fourth, sixth and eighth visual field. Like the other DL models, their model outperformed point-wise linear regression.
Yousefi et al. (36) showed Gaussian mixture-model with expectation maximisation (GEM) and variational Bayesian Independent Component Analysis mixture model (vB-ICAmm) were superior in detecting progression in glaucoma visual fields than PLR, permutation analysis of PLR, and linear regression using mean deviation. They investigated the detection of longitudinal progression in visual fields by clustering the visual fields into three groups: normal; early glaucoma; or advanced glaucoma (assessed by the visual field mean deviation value). GEM was compared against global, region-wise [glaucoma hemifield test (GHT)] and point-wise indices. The outcome of that study confirmed that ML algorithms were more sensitive than all other methods. However again these and other (25) statistical methods did not provide a visual representation of what the future visual field would look like, rather they provided either a classification or a prediction of a value similar to global indices.
The importance of spatial differences in visual fields was shown by Keltner et al. (28) who characterised patterns of field loss by manually inspecting abnormalities in the four visual hemi-fields of patients in the Ocular Hypertension Treatment Study. Using 2,509 HFA 30-2 visual fields, the abnormalities were first divided into superior and inferior hemi-fields before being further classified into one of 17 mutually exclusive categories of observed visual field loss patterns (28). The exclusive categories included patterns of VF loss characterised by glaucoma as well as other ocular and neurological diseases. From the pattern of VF loss characterised, 57.7% of the samples were judged to be typically glaucomatous, with the most frequent classification categories including partial arcuate, paracentral and nasal step defects. Non-glaucomatous fields were found in 20.3% of the samples, 11.5% of samples were associated with neurological abnormalities characterised by widespread, central and total loss, partial hemianopia, vertical-step and quadrant defects. Recent advances in artificial intelligence have produced models with the ability to simultaneously learn salient feature representations while completing cluster assignment (29). The Keltner 17 patterns of field loss were substantiated recently with unsupervised statistical learning ‘archetypal analysis’ (29).
We discovered that there was a strong correlation between the visual field locations, with a large variance inflation factor found for all except for the blind spot (data not presented). This reduces the reliability of the coefficients in the linear regression model, which was our comparative baseline test. Two solutions to multi-collinearity are proposed by James et al. (35); dropping the problematic variables (which is not a viable option for VFH data) and combining the collinear variables together into a single predictor. The second solution would require some dimensionality reduction, which would perhaps also be a useful future investigation.
The results from these models would be different for specific diagnoses and with the addition of more detailed clinical information (e.g., intra-ocular pressure, treatments). Certainly, Wen et al. (16) found the inclusion of the patient’s age improved their model’s predictability.
Apart from having a range of diagnoses in our dataset, there were also normal patients; understanding ‘normal’ test variance and the influence of age would improve future model development. In particular, the addition of Humphrey machine record of reliability indices in future models is likely to improve the outcomes.
The temporal sequence of the testing also hampers analysis, as in real life the results are not strictly standardised. It might be possible to utilise clinical trial data for comparison, as those data tend to more standardised, but the benefit for model refinement remains unclear.
It is generally accepted that patients may have a 2–3 dB variability in the test performance, and this is more so with glaucoma patients. A poor baseline test reliability compounds the problem of variability rendering any prediction inaccurate. Relating the Jensen’s inequality to the point wise prediction error, Wen et al. (16) suggested that the theoretical pointwise MAE lower limit was 2.32 dB, suggesting our C3D CNN result for the 4th sequence was close to the limit of predictability. The results for the small number of out of sequence patients were encouraging, again all the DL models outperformed the linear regression, at around the same prediction error (Table 7). This suggests the overall model fit is good, applying well to ‘unseen’ data.
Table 7
| ‘Out of sequence’ prediction RMSE results | 4th VFH in the sequence RMSE, dB | 5th VFH in the sequence RMSE, dB | 6th VFH in the sequence RMSE, dB |
|---|---|---|---|
| Simple RNN | 2.264 | 2.709 | 2.836 |
| Simple LSTM | 2.181† | 2.577 | 2.841 |
| GRU | 2.336 | 2.741 | 2.948 |
| Bidirectional RNN | 2.520 | 2.802 | 2.812 |
| Bidirectional LSTM | 2.452 | 2.701 | 2.410† |
| Bidirectional GRU | 2.840 | 2.947 | 2.831 |
| ConvLSTM single layer | 2.639 | 2.439 | 2.786 |
| ConvLSTM double layer | 2.567 | 2.401† | 2.754 |
| Modified C3D CNN | 2.401 | 2.691 | 3.141 |
| Linear regression | 4.042 | 3.881 | 4.420 |
† indicates the lowest RMSE in the column. Mean RMSE for the 4th, 5th, and 6th test result in the sequence. This compares favourably with the test set prediction results (Table 5) suggesting the overall model fit is good, applying well to ‘unseen’ data. C3D, three-dimensional convolution; CNN, convolutional neural network; Conv, convolutional; GRU, gated recurrent unit; LSTM, long short-term memory; RMSE, root mean square error; RNN, recurrent neural network; VFH, Humphrey visual field.
Conclusions
Although these results confirm that overall DL models are better for predicting visual field loss than the presently used ordinary least squares regression model, we found no one model superior.
Whether the addition of spatial information to a temporal sequence model improves prediction, however, remains unclear. Certainly, we found the standard and commonly used computer-vision DL models to date did not greatly enhance predictability.
Whilst our findings and those of others haven’t found an optimal solution for spatio-temporal prediction of visual field loss progression, future work exploring the spatial characteristics of visual field loss such as archetypal analysis may provide fruitful results.
Acknowledgments
None.
Footnote
Reporting Checklist: The authors have completed the TRIPOD+AI reporting checklist. Available at https://aes.amegroups.com/article/view/10.21037/aes-25-38/rc
Data Sharing Statement: Available at https://aes.amegroups.com/article/view/10.21037/aes-25-38/dss
Peer Review File: Available at https://aes.amegroups.com/article/view/10.21037/aes-25-38/prf
Funding: None.
Conflicts of Interest: All authors have completed the ICMJE uniform disclosure form (available at https://aes.amegroups.com/article/view/10.21037/aes-25-38/coif). J.N. is the committee member of Australian and New Zealand Glaucoma Society and received grants from Australian Research Council. The other authors have no conflicts of interest to declare.
Ethical Statement:
Open Access Statement: This is an Open Access article distributed in accordance with the Creative Commons Attribution-NonCommercial-NoDerivs 4.0 International License (CC BY-NC-ND 4.0), which permits the non-commercial replication and distribution of the article with the strict proviso that no changes or edits are made and the original work is properly cited (including links to both the formal publication through the relevant DOI and the license). See: https://creativecommons.org/licenses/by-nc-nd/4.0/.
References
- Heijl A, Patella VM. Essential Perimetry. 3rd ed. Jena: Carl Zeiss Meditek; 2002.
- De Moraes CG, Liebmann JM, Levin LA. Detection and measurement of clinically meaningful visual field progression in clinical trials for glaucoma. Prog Retin Eye Res 2017;56:107-47. [Crossref] [PubMed]
- European Glaucoma Society Terminology and Guidelines for Glaucoma, 4th Edition - Part 1 Supported by the EGS Foundation. Br J Ophthalmol 2017;101:1-72.
- Heijl A, Bengtsson B. The effect of perimetric experience in patients with glaucoma. Arch Ophthalmol 1996;114:19-22. [Crossref] [PubMed]
- Heijl A, Lindgren G, Olsson J. The effect of perimetric experience in normal subjects. Arch Ophthalmol 1989;107:81-6. [Crossref] [PubMed]
- Wild JM, Dengler-Harles M, Searle AE, et al. The influence of the learning effect on automated perimetry in patients with suspected glaucoma. Acta Ophthalmol (Copenh) 1989;67:537-45. [Crossref] [PubMed]
- Wild JM, Searle AE, Dengler-Harles M, et al. Long-term follow-up of baseline learning and fatigue effects in the automated perimetry of glaucoma and ocular hypertensive patients. Acta Ophthalmol (Copenh) 1991;69:210-6. [Crossref] [PubMed]
- Leske MC, Heijl A, Hyman L, et al. Early Manifest Glaucoma Trial: design and baseline data. Ophthalmology 1999;106:2144-53. [Crossref] [PubMed]
- Tanna AP, Budenz DL, Bandi J, et al. Glaucoma Progression Analysis software compared with expert consensus opinion in the detection of visual field progression in glaucoma. Ophthalmology 2012;119:468-73. [Crossref] [PubMed]
- Heijl A, Patella V, Bengtsson B. Effective perimetry: The field analyzer primer. 4th ed. Jena: Carl Zeiss Meditek; 2012.
- Goldbaum MH. Unsupervised learning with independent component analysis can identify patterns of glaucomatous visual field defects. Trans Am Ophthalmol Soc 2005;103:270-80. [PubMed]
- Goldbaum MH, Lee I, Jang G, et al. Progression of patterns (POP): a machine classifier algorithm to identify glaucoma progression in visual fields. Invest Ophthalmol Vis Sci 2012;53:6557-67. [Crossref] [PubMed]
- Sample PA, Chan K, Boden C, et al. Using unsupervised learning with variational bayesian mixture of factor analysis to identify patterns of glaucomatous visual field defects. Invest Ophthalmol Vis Sci 2004;45:2596-605. [Crossref] [PubMed]
- Berchuck SI, Mukherjee S, Medeiros FA. Estimating Rates of Progression and Predicting Future Visual Fields in Glaucoma Using a Deep Variational Autoencoder. Sci Rep 2019;9:18113. [Crossref] [PubMed]
- Park K, Kim J, Lee J. Visual Field Prediction using Recurrent Neural Network. Sci Rep 2019;9:8385. [Crossref] [PubMed]
- Wen JC, Lee CS, Keane PA, et al. Forecasting future Humphrey Visual Fields using deep learning. PLoS One 2019;14:e0214875. [Crossref] [PubMed]
- Eslami M, Kim JA, Zhang M, et al. Visual Field Prediction: Evaluating the Clinical Relevance of Deep Learning Models. Ophthalmol Sci 2022;3:100222. [Crossref] [PubMed]
- Kim H, Lee J, Moon S, et al. Visual field prediction using a deep bidirectional gated recurrent unit network model. Sci Rep 2023;13:11154. [Crossref] [PubMed]
- Marquez ES, Hare JS, Niranjan M. Deep Cascade Learning. IEEE Trans Neural Netw Learn Syst 2018;29:5475-85. [Crossref] [PubMed]
- Rumelhart DE, Hinton GE, Williams RJ. Learning internal representations by error propagation. In: Parallel Distributed Processing: Explorations in the Microstructure of Cognition. Cambridge, MA, USA: MIT Press; 1986:318-62.
- Hochreiter S, Schmidhuber J. Long short-term memory. Neural Comput 1997;9:1735-80. [Crossref] [PubMed]
- Cho K, van Merrienboer B, Gulcehre C, et al. Learning phrase representations using RNN encoder-decoder for statistical machine translation. arXiv:1406.1078 [Preprint]. 2014. Available online: https://doi.org/10.48550/arXiv.1406.1078.
- LeCun Y, Haffner P, Bottou L, et al. Object recognition with gradient-based learning. In: Shape, contour and grouping in computer vision. Berlin, Heidelberg: Springer; 1999:319-45.
- Shi X, Chen Z, Wang H, et al. Convolutional LSTM network: A machine learning approach for precipitation nowcasting. arXiv:1506.04214 [Preprint]. 2015. Available online: https://arxiv.org/abs/1506.04214.
- Fujino Y, Murata H, Mayama C, et al. Applying "Lasso" Regression to Predict Future Visual Field Progression in Glaucoma Patients. Invest Ophthalmol Vis Sci 2015;56:2334-9. [Crossref] [PubMed]
- Wang M, Shen LQ, Boland MV, et al. Impact of Natural Blind Spot Location on Perimetry. Sci Rep 2017;7:6143. [Crossref] [PubMed]
- Chai T, Draxler RR. Root mean square error (RMSE) or mean absolute error (MAE)? – Arguments against avoiding RMSE in the literature. Geosci Model Dev 2014;7:1247-50.
- Keltner JL, Johnson CA, Cello KE, et al. Classification of visual field abnormalities in the ocular hypertension treatment study. Arch Ophthalmol 2003;121:643-50. [Crossref] [PubMed]
- Elze T, Pasquale LR, Shen LQ, et al. Patterns of functional vision loss in glaucoma determined with archetypal analysis. J R Soc Interface 2015;12:20141118. [Crossref] [PubMed]
- Berchuck SI, Medeiros FA, Mukherjee S. Scalable modelling of Spatio-temporal data using the variational auto-encoder: an application in glaucoma. arXiv:1908.09195 [Preprint]. 2019. Available online: https://arxiv.org/abs/1908.09195.
- Tran D, Bourdev L, Fergus R, et al. Learning spatio-temporal features with 3d convolutional networks. 2015 IEEE International Conference on Computer Vision (ICCV); Santiago, Chile. IEEE; 2015:4489-97.
- Geron A. Hands-On Machine Learning with Scikit-Learn, Keras, and TensorFlow: Concepts, Tools, and Techniques to Build Intelligent Systems. 1st ed. Sebastopol, CA: O'Reilly Media; 2022.
- Hyperparameter Experiments with TensorFlow and Keras. Available online: https://github.com/autonomio/talos.
- Srivastava N, Hinton G, Krizhevsky A, et al. Dropout: a simple way to prevent neural networks from overfitting. The Journal of Machine Learning Research 2014;15:1929-58.
- James G, Witten D, Hastie T, et al. An Introduction to Statistical Learning: with Applications in R. New York, NY: Springer; 2013.
- Yousefi S, Kiwaki T, Zheng Y, et al. Detection of Longitudinal Visual Field Progression in Glaucoma Using Machine Learning. Am J Ophthalmol 2018;193:71-9. [Crossref] [PubMed]
Cite this article as: Hu ML, Morlet N, Liu W, Glance D, Mutiawani V, Morgan B, Manners S, Ng J. Deep learning artificial intelligence models compared to ordinary linear regression for prediction of visual field progression. Ann Eye Sci 2026;11:23.

