Confirmatory Factor Analysis

Exploratory factor analysis can be used to identify common factors and factor structure among a set of observed variables / indicators. Confirmatory factor analysis (CFA) can be used to study how well a hypothesized factor model fits a new sample from the same population or a sample from a different population. The CFA model is the same as the EFA model with the exception that restrictions can be placed on factor loadings, variances, covariances, and residual variances resulting in a more parsimonious model. Using CFA, one can

  • investigate if a factor model fits a new sample from the same population – the confirmatory aspect.
  • evaluate if a factor model fits a sample from a different population – measurement invariance,
  • study the behavior of new measurement items embedded in a previously studied measurement instrument, and
  • estimate factor scores.

An example

The study by Karl J. Holzinger and Frances Swineford involves data collection from two schools, the Grant-White School (\(n=145\)) and Pasteur School (\(n=156\)). In the EFA example, we have identified 4 factors with data from the the Grant-White School. We now investigate whether the same factor model would fit the data in the Pasteur School. The data are saved in the file Pasteur.csv. 

Confirmatory factor analysis

The model

The path diagram for the model is given in the figure below. Note that we have assumed there are 4 factors. And the indicators of each factor are also known. This model is based on the EFA of Grant-White school data with the factor loadings greater than 0.3 kept in the model.

Model estimation

To estimate a confirmatory factor model, the R package lavaan can used. A confirmatory factor model cannot be identified without proper constraints, that's, to fix some parameters to be known values in the model. The reason is that factors are unmeasured and thus have no scales. To identify a model, the factors have to be given specific scales. There are two ways to do this. First, we can fix the variance of a factor to be 1. This is to standardize the factor. Second, we can fix the loading of one observed variable or indicator to be one. This is essentially to make the factor to have the same scale as the observed variables. Fixing a factor loading or a factor variance is the minimum requirement to identify a factor model. However, this does not guarantee a model is identifiable. Practically, if a model cannot be identified, the software used to estimate the model will not run correctly. 

The cfa() function in lavaan can be used to estimate a factor model. To use the function, we need to first specify the factor model. A factor model take the format

factor =~ y1 + y2 +y3

Note that we use the symbol "=~" to define a factor. The factor is on the left of the symbol and the indicators are on the right of it.

The R code for the example is given below. 

> library(lavaan)
This is lavaan 0.5-23.1097
lavaan is BETA software! Please report any bugs.
> usedata('Pasteur.csv')
> 
> cfa.model<-'
+ spatial =~ visual+cubes+paper+lozenge+straight+figurer
+ verbal =~ general+paragrap+sentence+wordc+wordm
+ speed =~ add+code+counting+straight
+ memory =~ wordr+numberr+figurer+object+numberf+figurew
+ '
> 
> cfa.est<-cfa(cfa.model, data=Pasteur)
> summary(cfa.est, fit=TRUE)
lavaan (0.5-23.1097) converged normally after 251 iterations

  Number of observations                           156

  Estimator                                         ML
  Minimum Function Test Statistic              238.751
  Degrees of freedom                               144
  P-value (Chi-square)                           0.000

Model test baseline model:

  Minimum Function Test Statistic             1149.278
  Degrees of freedom                               171
  P-value                                        0.000

User model versus baseline model:

  Comparative Fit Index (CFI)                    0.903
  Tucker-Lewis Index (TLI)                       0.885

Loglikelihood and Information Criteria:

  Loglikelihood user model (H0)              -9899.892
  Loglikelihood unrestricted model (H1)      -9780.517

  Number of free parameters                         46
  Akaike (AIC)                               19891.785
  Bayesian (BIC)                             20032.078
  Sample-size adjusted Bayesian (BIC)        19886.474

Root Mean Square Error of Approximation:

  RMSEA                                          0.065
  90 Percent Confidence Interval          0.050  0.079
  P-value RMSEA <= 0.05                          0.050

Standardized Root Mean Square Residual:

  SRMR                                           0.079

Parameter Estimates:

  Information                                 Expected
  Standard Errors                             Standard

Latent Variables:
                   Estimate  Std.Err  z-value  P(>|z|)
  spatial =~                                          
    visual            1.000                           
    cubes             0.396    0.089    4.465    0.000
    paper             0.233    0.051    4.547    0.000
    lozenge           1.087    0.182    5.961    0.000
    straight          1.599    0.623    2.567    0.010
    figurer           0.535    0.139    3.836    0.000
  verbal =~                                           
    general           1.000                           
    paragrap          0.282    0.024   11.846    0.000
    sentence          0.463    0.035   13.322    0.000
    wordc             0.388    0.038   10.177    0.000
    wordm             0.589    0.047   12.585    0.000
  speed =~                                            
    add               1.000                           
    code              0.835    0.145    5.758    0.000
    counting          0.708    0.147    4.811    0.000
    straight          1.091    0.265    4.115    0.000
  memory =~                                           
    wordr             1.000                           
    numberr           0.537    0.106    5.040    0.000
    figurer           0.486    0.106    4.570    0.000
    object            0.379    0.070    5.372    0.000
    numberf           0.277    0.060    4.638    0.000
    figurew           0.245    0.055    4.444    0.000

Covariances:
                   Estimate  Std.Err  z-value  P(>|z|)
  spatial ~~                                          
    verbal           21.927    5.770    3.800    0.000
    speed            24.531    9.382    2.615    0.009
    memory           10.542    4.998    2.109    0.035
  verbal ~~                                           
    speed            66.093   17.349    3.810    0.000
    memory           10.027    7.727    1.298    0.194
  speed ~~                                            
    memory           47.001   14.950    3.144    0.002

Variances:
                   Estimate  Std.Err  z-value  P(>|z|)
   .visual           21.376    4.533    4.716    0.000
   .cubes            19.547    2.401    8.139    0.000
   .paper             6.472    0.799    8.102    0.000
   .lozenge          52.062    7.715    6.748    0.000
   .straight        875.004  110.732    7.902    0.000
   .figurer          40.094    5.568    7.200    0.000
   .general          41.308    5.979    6.909    0.000
   .paragrap          4.185    0.571    7.330    0.000
   .sentence          6.615    1.058    6.252    0.000
   .wordc            13.215    1.664    7.943    0.000
   .wordm            14.218    2.063    6.892    0.000
   .add             415.730   55.840    7.445    0.000
   .code             75.538   19.745    3.826    0.000
   .counting        280.324   35.795    7.831    0.000
   .wordr            83.348   12.917    6.452    0.000
   .numberr          43.595    5.790    7.529    0.000
   .object           16.578    2.329    7.119    0.000
   .numberf          15.587    1.982    7.865    0.000
   .figurew          14.000    1.752    7.989    0.000
    spatial          28.852    6.417    4.496    0.000
    verbal           96.582   15.354    6.290    0.000
    speed           200.381   59.608    3.362    0.001
    memory           60.814   15.974    3.807    0.000

> 

Output

By default, the cfa() function fixes a factor loading for each factor to be 1 and estimates the rest factor loadings. Different from EFA, CFA also test the significance of the factor loadings based on a z-test. For example, the factor loading for cubes on factor 1 is 0.396 with the standard error 0.089. The corresponding z-value is 4.465 with a p-value almost 0. Therefore, this factor loading is statistically significant from 0. In addition, it also estimates the factor variances and covariances as well as the uniqueness factor variances. 

Model fit evaluation

CFA provides a lot more information to evaluate whether a model fits a sample. First, a chi-square test can be used. For this test, the null hypothesis is that the model fits the data well. Therefore, one would hope to get a small chi-square statistic and a corresponding large p-value. Typically, when p-value is larger than 0.5, one fails to reject the model.

There are many other fit statistics and indices that can be used to evaluate model fit. The most widely used ones include CFI, RMSEA, and SRMR. 

  • Comparative fit index (CFI) compares the model under evaluation with a baseline model. One form of the baseline model is the independent model where it assumes the independence among the observed variables. In general, such a baseline model would not fit the data well. Therefore, CFI measures the relative improvement of the current model. It is generally accepted that a CFI greater than 0.96 (or 0.95) indicates a good fit model.
  • Root mean square error of approximation (RMSEA) is a measure of chi-square attributed to each participate after controlling the model complexity. Therefore, a smaller value indicates better fit. In the literature, 
    • a RMSEA $\leq 0.05$ indicates a model fits the data closely. 
    • a $0.05<$ RMSEA $\leq 0.08$ indicates a model fits the data reasonably well.
    • a RMSEA $>0.1$ indicates a model is a bad model.
  • Standardized root mean square residual (SRMR) measures the difference between the observed covariance matrix from the data and the predicted covariance matrix based on the model. A value SRMR=0 indicates a perfect of the model. In general, a value less than 0.08 is considered a good fit of the model under evaluation.

Using these criteria, we can evaluate whether the confirmatory factor model identified using the data from Grant-White school fits the data from the Pasteur school well. 

  • The chi-square statistic (Minimum Function Test Statistic in the output) is 238.75 with the degrees of freedom 144 and a p-value close to 0. Therefore, one would reject the hypothesis that the model fits the data simply based on it.
  • Comparative Fit Index (CFI) is 0.903, which is smaller than the cut-off value 0.95. It also suggests a bad fit.
  • The RMSEA = 0.065, which lies the range of a reasonable fit model.
  • The SRMR = 0.079, which is smaller than but close to the cut-off value 0.08.

Overall, the chi-square test and the other criteria suggest the model barely fits the data. Therefore, we cannot replicate the factor structure identified in EFA for the Grant-White school data in the Pasteur school.

Estimate the model with standardized factors

Instead of fitting a factor loading to be 1, we can also fix the factor variances to be 1. The R code below does the analysis. Note that the model fit is exactly the same as before. However, we can now directly estimate the factor correlation matrix.

> library(lavaan)
This is lavaan 0.5-23.1097
lavaan is BETA software! Please report any bugs.
> usedata('Pasteur.csv')
> 
> cfa.model<-'
+ spatial =~ visual+cubes+paper+lozenge+straight+figurer
+ verbal =~ general+paragrap+sentence+wordc+wordm
+ speed =~ add+code+counting+straight
+ memory =~ wordr+numberr+figurer+object+numberf+figurew
+ '
> 
> cfa.est<-cfa(cfa.model, data=Pasteur, std.lv=TRUE)
> summary(cfa.est, fit=TRUE)
lavaan (0.5-23.1097) converged normally after 191 iterations

  Number of observations                           156

  Estimator                                         ML
  Minimum Function Test Statistic              238.751
  Degrees of freedom                               144
  P-value (Chi-square)                           0.000

Model test baseline model:

  Minimum Function Test Statistic             1149.278
  Degrees of freedom                               171
  P-value                                        0.000

User model versus baseline model:

  Comparative Fit Index (CFI)                    0.903
  Tucker-Lewis Index (TLI)                       0.885

Loglikelihood and Information Criteria:

  Loglikelihood user model (H0)              -9899.892
  Loglikelihood unrestricted model (H1)      -9780.517

  Number of free parameters                         46
  Akaike (AIC)                               19891.785
  Bayesian (BIC)                             20032.078
  Sample-size adjusted Bayesian (BIC)        19886.474

Root Mean Square Error of Approximation:

  RMSEA                                          0.065
  90 Percent Confidence Interval          0.050  0.079
  P-value RMSEA <= 0.05                          0.050

Standardized Root Mean Square Residual:

  SRMR                                           0.079

Parameter Estimates:

  Information                                 Expected
  Standard Errors                             Standard

Latent Variables:
                   Estimate  Std.Err  z-value  P(>|z|)
  spatial =~                                          
    visual            5.371    0.597    8.993    0.000
    cubes             2.124    0.433    4.901    0.000
    paper             1.254    0.250    5.012    0.000
    lozenge           5.837    0.790    7.385    0.000
    straight          8.589    3.240    2.651    0.008
    figurer           2.872    0.697    4.121    0.000
  verbal =~                                           
    general           9.828    0.781   12.580    0.000
    paragrap          2.773    0.234   11.850    0.000
    sentence          4.551    0.340   13.389    0.000
    wordc             3.811    0.375   10.172    0.000
    wordm             5.790    0.459   12.605    0.000
  speed =~                                            
    add              14.156    2.105    6.723    0.000
    code             11.823    1.224    9.658    0.000
    counting         10.024    1.673    5.990    0.000
    straight         15.440    3.235    4.773    0.000
  memory =~                                           
    wordr             7.798    1.024    7.614    0.000
    numberr           4.185    0.683    6.129    0.000
    figurer           3.790    0.707    5.363    0.000
    object            2.953    0.434    6.798    0.000
    numberf           2.160    0.398    5.428    0.000
    figurew           1.913    0.374    5.121    0.000

Covariances:
                   Estimate  Std.Err  z-value  P(>|z|)
  spatial ~~                                          
    verbal            0.415    0.085    4.895    0.000
    speed             0.323    0.103    3.129    0.002
    memory            0.252    0.109    2.303    0.021
  verbal ~~                                           
    speed             0.475    0.081    5.902    0.000
    memory            0.131    0.098    1.336    0.181
  speed ~~                                            
    memory            0.426    0.097    4.410    0.000

Variances:
                   Estimate  Std.Err  z-value  P(>|z|)
   .visual           21.376    4.533    4.716    0.000
   .cubes            19.547    2.401    8.139    0.000
   .paper             6.472    0.799    8.102    0.000
   .lozenge          52.062    7.715    6.748    0.000
   .straight        875.006  110.732    7.902    0.000
   .figurer          40.093    5.568    7.200    0.000
   .general          41.308    5.979    6.909    0.000
   .paragrap          4.185    0.571    7.330    0.000
   .sentence          6.615    1.058    6.252    0.000
   .wordc            13.215    1.664    7.943    0.000
   .wordm            14.218    2.063    6.892    0.000
   .add             415.731   55.841    7.445    0.000
   .code             75.538   19.745    3.826    0.000
   .counting        280.324   35.795    7.831    0.000
   .wordr            83.347   12.917    6.452    0.000
   .numberr          43.595    5.790    7.529    0.000
   .object           16.578    2.329    7.119    0.000
   .numberf          15.587    1.982    7.865    0.000
   .figurew          14.000    1.752    7.989    0.000
    spatial           1.000                           
    verbal            1.000                           
    speed             1.000                           
    memory            1.000                           

> 

To cite the book, use: Zhang, Z. & Wang, L. (2017-2026). Advanced statistics using R. Granger, IN: ISDSA Press. https://doi.org/10.35566/advstats. ISBN: 978-1-946728-01-2.
To take the full advantage of the book such as running analysis within your web browser, please subscribe.