
phylosem model description
James T. Thorson
Source:vignettes/model-description.Rmd
model-description.RmdPhylogenetic structural equation models
phylosem is an R package for fitting phylogenetic
structural equation models (PSEMs). The package generalizes features in
existing R packages:
-
semfor structural equation models (SEMs); -
phylosemfor comparison among alternative path models; -
phylolmfor fitting large linear models that arise as when specifying a SEM with one endogenous variable and multiple exogenous and independent variables; -
Rphyloparsfor interpolating missing values when specifying a SEM with an unstructured (full rank) covariance among variables;
In model configurations that can be fitted by both
phylosem and these other packages, we have confirmed that
results are nearly identical or otherwise identified reasons that
results differ.
phylosem involves a simple user-interface that specifies
the SEM using notation from package sem and the
phylogenetic tree using package ape. It allows uers to
specify common models for the covariance including:
- Brownian motion (BM);
- Ornstein-Uhlenbeck (OU);
- Pagelβs lambda;
- Pagelβs kappa;
Output can be coerced to standard formats so that
phylosem can use plotting and summary functions form other
packages. Available output formats include:
-
sem, for plotting the estimated SEM and summarizing direct and indirect effects; -
phylopath, for plotting and model comparison; -
phylo4din R-packagephylobasefor plotting estimated traits;
Package phylosem (Thorson and Bijl 2023) involves specifying a phylogenetic structural equation model (PSEM). This PSEM be viewed either:
- Weak interpretation: as an expressive interface to parameterize the correlation among life-history traits, using as many or few parameters as might be appropriate; or
- Strong interpretation: as a structural causal model, allowing predictions about the consequence of counterfactual changes to the system.
We introduce PSEM first from the perspective of a software user (i.e., the interface) and then from the perspective of a statistician (i.e., the equations and their interpretation).
Viewpoint 1: Software interface
To specify a PSEM, the user uses arrow notationderived from
package sem (Fox 2006). For
example, to specify that evolutionary chnages in trait
causes a change in trait
,
we specify:
This then estimates a single parameter representing the impact of
on
(specified with one-headed arrows), as well as the square root (Cholesky
decomposition) of the exogenous covariance of of model variables
(specified with two-headed arrows). See ?phylosem Details
section for more details about syntax.
PSEM interactions can be as complicated or simple as desired, and can include:
- Latent variables and loops (i.e., they are not restricted to directed acyclic graphs);
- Values that are fixed a priori, where the
parameter_nameis provided asNAand the starting value that follows is the fixed value; - Values that are mirrored among path coefficients, where the same
parameter_nameis provided for multiple rows of the text file.
The user also specifies a distribution for measurement errors for
each variable using arguement family, and a dated
phylogenetic tree tree that is used to represent
evolutionary correlations.
Viewpoint 2: Mathematical details
The PSEM defines a generalized linear mixed model (GLMM) for a matrix , where is the measurement for vertex for variable , where vertex can be either an ancestral node or a tip (i.e., PSEM estimates ancestral traits jointly with tips). This measurement matrix can include missing values , and it will estimate a matrix of latent states for all modeled vertices and variables, where is the vector of traits for vertex and is a length vector constructed by stacking columns. PSEM also estimates a matrix of process errors where is the vector of process errors for vertex , and is the vector of stacked columns.
PSEM can be written in structural vector-autoregressive (SVAR) notation as a lag-1 SVAR model :
where this expression is evaluated for each edge of the phylogenetic tree connecting child with parent . is the matrix of relationships among traits, and is exogenous covariation for vertex :
and is the square-root of exogenous covariance that occurs in each time. is then estimated by PSEM, and is identifiable without constraints when specified as a lower-triangle matrix (representing the Cholesky decomposition of the covariance of exogenous process errors occurring in each time).
This SVAR notation can also be rewritten as a joint simultaneous equation model (SEM), although we do not elaborate here ((thorson_graphical_2026?))
Measurement errors
PSEM includes multiple distribution for measurement errors. For
example, if the user specifies family[[j] = fixed()
then:
for all vertices, where
is an estimated mean parameter that is only estimated when specifying an
Ornstein-Uhlenbeck process (and otherwise fixed at zero). Alternatively,
if the user specifies family[[j]] = gaussian() then:
and
is then included as an estimated parameter. When estimating missing
values or masurement errors in
,
phylosem must then marginalize across the latent value of
states
.
It does this using the Laplace approximation (Skaug and Fournier 2006), as implemented using
the R-package TMB (Kristensen et al.
2016). Computations involving sparse matrices are efficient using
the Matrix package (Bates et al. 2023) to
interface with the Eigen library (Guennebaud et al. 2010).