Skip to main content
Have a personal or library account? Click to login
CodeEntropy: A Python Package for Multiscale Cell Correlation Entropy Estimation from Molecular Dynamics Simulations Cover

CodeEntropy: A Python Package for Multiscale Cell Correlation Entropy Estimation from Molecular Dynamics Simulations

Open Access
|Jul 2026

Full Article

(1) Overview

Introduction

Entropy is a fundamental thermodynamic quantity of molecular systems, yet remains one of the most challenging to compute accurately. This is because its evaluation involves determining the probabilities of all quantum states of the system

1
S=kBΣiρilnρi

where kB is Boltzmann’s constant and ρi is the probability of each quantum state. In general, this is an intractable problem to tackle directly owing to the enormous number of states involved. To make any progress requires a careful set of approximations to retain accuracy and enable practicality. The multiscale cell correlation (MCC) method offers a general and unifying framework for estimating entropy from trajectories of atom forces and coordinates generated in a molecular dynamics (MD) simulation. The advantages of MCC over other entropy estimators is its consistent and parameter-free treatment of all molecules that enables generality and automatability, its quantum nature over all atomic degrees of freedom that produces absolute entropy, and its simple and hierarchical formulation that provides speed, scalability, and intuitive interpretation. Regarding MCC’s limitations, it uses the quantum harmonic approximation derived from a classical MD simulation and force-field, it accounts for only a subset of correlations, and it adopts a coarse theory for the discretisation of energy minima. MCC has been developed, tested, and generalised by Henchman and co-workers on a range of progressively more complex systems, ranging from liquids (1, 2, 3, 4) and small-molecule solutions (5, 6, 7, 8, 9, 10) to proteins (11, 12, 13) and complexes (14, 15, 16). Each system involved the development of in-house software that contained new theory specific to the systems under consideration and was not yet suitably purposed for general use. CodeEntropy addresses this gap by providing an open-source, well-engineered, modular and general implementation of MCC that integrates the generalised theory in an extensible software architecture under one roof.

MCC utilises atom coordinates and forces from a single equilibrium MD simulation of the system, quantities that are output by most simulation codes. MCC’s multiscale nature arises from grouping molecules into progressively larger structural sets of atoms, termed beads, at different levels of structure. From smallest to largest, beads can comprise atoms, united atoms, residues, polymers or assemblies. Nuclear and electronic degrees of freedom are not considered, and nor are those of higher-order structures. A level with only one bead is by definition a molecule. Each bead has a local coordinate system that accurately captures flexibility at every length scale. If beads are weakly correlated with other beads at the same level, then they experience mean-field cell potentials, which occur for beads that are molecules or for bead rotation. If beads are strongly correlated, they experience orthogonal harmonic potentials for collective modes over multiple atoms, as occurs for the translation of covalently bonded beads.

Theoretical Background

At the heart of MCC is the separation of the full probability distribution of all quantum states of the system into a product of lower-dimensional distributions of beads at each length scale involving translation and rotation in localised Cartesian coordinates; these distributions are further discretised into a configurational term over multiple minima and a vibrational term averaged over these minima. Total system entropy is therefore evaluated as a sum of terms Smlovib+Smloconfig over molecule type m, molecule level l, and motion type o either translation (t) or rotation (r) expressed by

2
S=ΣMm=1NmΣLml=1Σo=t,r(Smlovib+Smloconfig)

where M is the number of molecule types, Nm is the number of molecules of type m, and Lm is the number of levels for molecule type m. CodeEntropy currently implements the united-atom, residue, and polymer levels of hierarchy.

Vibrational Entropy

The translational and rotational components of vibrational entropy of a set of beads at a given level l and given type of molecule m are calculated in the harmonic approximation using the quantum harmonic oscillator equation

3
Smlovib=kBΣi=1Nmlovib(hνmloivib/kBTehνmloivib/kBT1ln(1ehνmloivib/kBT))

where Nmlovib is the number of coordinates for all beads at level mlo, h is Planck’s constant, T is temperature and νmloivib are the characteristic frequencies of each vibrational eigenvector. The frequencies are derived from the eigenvalues λmloivib of the covariance matrix of mass-weighted forces for translation and of inertia-weighted torques for rotation via

4
νmloivib=12πλmloivibkBT

To construct each covariance matrix, forces and torques on each bead at every time step are transformed into the local coordinate frame of each bead. The axes at the molecule level are taken as the principal axes with origin at the centre of mass. Translation of beads within a molecule at the level below uses the same molecular axes, while rotation uses local bead axes that are constructed from other beads covalently bonded to it, with origin at the central bead’s position (3, 12). This pattern continues if there is an even lower level. The entropy of the six lowest-frequency translational eigenvectors is excluded to avoid double-counting entropy at the higher level of hierarchy. The vibrational term is omitted at the atomic level because of almost complete occupancy in the ground quantum state at ambient conditions, owing to hydrogen’s light mass and high-frequency bond vibration making a high-energy gap to the first excited state. Forces and torques are halved for degrees of freedom treated with the mean-field approximation, so as to partition all pairwise forces equally between the two beads contributing to that force. In practice, this means that torques are halved at all levels and forces are halved at the molecule level but not for internal translation due to correlation across covalent bonds. The exception to this is molecules with flexible dihedrals, as defined in the next section, whose flexibility makes them decorrelated. Consequently, the eigenvalues of the lowest Nmlodih eigenvectors, corresponding to these soft modes, are halved twice (4), where Nmlodih is the number of flexible dihedrals.

Configurational Entropy

The configurational term quantifies the entropy associated with the probability distribution pmlojconfig of energy well j at level mlo and is evaluated using Eq. 1. Configurational entropy comprises three kinds of term: conformational, orientational, and positional entropy. These reflect different levels of hierarchy and whether it involves translation or rotation. Previous works (3, 4, 9, 10, 12, 13, 15, 16) had referred to these terms as “topographical.”

Conformational entropy involves translation of intramolecular beads. To calculate it, dihedrals are computed using the MDAnalysis package (17, 18) for all quadruplets of contiguously bonded beads. At the united atom level, dihedrals are defined from four bonded heavy atoms within a residue, while at the residue level, dihedrals are defined using the connecting bond vectors of four consecutive residues. The resulting dihedral distributions are binned into coarse histograms and distinct conformers are assigned to each peak in the histogram (12). Each dihedral j in the bead is assigned to the nearest state, allowing calculation of the probability pmljconform for the set of dihedrals in the bead, noting that “conform” represents both translation and configuration at level l. Entropy is then evaluated using Eq. 1. This procedure has been applied at both the united-atom and residue levels. Configurational entropy for rotation of united atoms is currently not included in CodeEntropy owing to its small contribution because of symmetry or correlation with neighbouring molecules.

Orientational entropy Smorient of molecule m arises from the discretisation of a molecule’s rotation by neighbouring molecules in its environment, where we note that the superscript “orient” represents rotation and configuration. The number of orientations Ωorient relates to the number of molecular neighbours Nmnb which is calculated using the relative angular distance (RAD) method (19). This parameter-free method defines a molecule j to be in the coordination shell of the central molecule i if for all other molecules k

5
1rij2>1rik2cosθjik

where, rij is the distance between i and j and θjik is the angle between j, i, and k (with i at the vertex). The number of orientations Ωmorient for non-linear molecules, which have three rotational degrees of freedom, is given by

6
Ωmorient=max{1,(Nmnb)3/2π1/2/σm}

and for linear molecules, which have two rotational degrees of freedom, this number is

7
Ωmorient=max{1,Nmnb/σm}

where σm is the symmetry number of the molecule, Nmnb is the average number of molecular neighbours, and the max function ensures there is at least one orientation. The resulting entropy

8
Smorient=kln(Ωmorient)

is equivalent to Eq. 1 with probability 1/Ωmorient. RDKit is used to determine σ using the GetSubstructMatches function and to determine if the molecule is linear using the GetHybridization function. Molecules with two united atom beads are treated as linear. If a molecule has N+3 united atom beads and at least N–2 of the heavy atoms have sp hybridization, then the molecule is considered linear for this orientational entropy calculation. CodeEntropy includes a second method to evaluate Smorient for water specifically as the solvent in the software waterEntropy that accounts for hydrogen-bond correlations at the atom level between neighbouring molecules and is discussed elsewhere (15, 13).

Positional entropy, commonly referred to as the entropy of mixing, arises from the discretisation of a molecule’s translation by the other molecules in the system. This term has not yet been included in CodeEntropy, but is included in future plans.

Implementation and Architecture

CodeEntropy is an open-source Python (>=3.12) package that implements the MCC method for estimating entropy from MD simulations. The software provides a modular analysis pipeline that constructs molecular hierarchies and accumulates statistical observables from trajectory data. The implementation prioritises methodological transparency, numerical reproducibility, and traceability while maintaining sufficient performance for analysing long MD trajectories. Entropy estimators are implemented using well-understood statistical and linear algebra operations, allowing calculations to remain auditable and scientifically interpretable.

Configuration-driven workflows and structured output formats ensure that analyses can be reproduced exactly and examined programmatically. Analysis parameters may be supplied through YAML (YAML Ain’t Markup Language) configuration files and selectively overridden through command-line arguments, enabling reproducible execution without modification of the source code.

The implementation relies on standard scientific Python libraries. Core numerical operations use NumPy (20) for array operations, statistical accumulation, and linear algebra. Graph construction and dependency resolution for the analysis workflow are implemented using NetworkX (21), which provides the representation of the directed acyclic graph (DAG) structures used to organise analysis nodes and determine execution order. Molecular trajectory parsing and atom selection are handled through MDAnalysis (17), while interactive terminal summaries are rendered using rich. Optional frame-parallel execution is provided through a Dask-based execution layer (22), allowing frame-local work to be distributed while preserving the same accumulated statistical state used by the serial workflow. This dependency stack keeps the implementation aligned with the established Python ecosystem for molecular simulation analysis.

Software contribution

The primary software contribution of CodeEntropy is the representation of MCC analysis as an explicit dependency-driven workflow that maps theoretical transformations directly onto executable computational stages. The implementation separates workflow orchestration, molecular hierarchy construction, frame-local observable evaluation, streaming statistical accumulation, conformational state construction, and final entropy estimation into modular components.

This design exposes intermediate quantities as first-class objects within the workflow, enabling inspection, validation, and extension of individual analytical stages. It also separates frame-independent molecular topology from frame-dependent numerical calculations, allowing static information such as hierarchy definitions, bead mappings, conformational topology, and coordinate-frame topology to be constructed once and reused throughout trajectory processing.

Motivation and design drivers

CodeEntropy is intended as an open and inspectable research implementation of the MCC framework. Its architecture is designed to reflect the conceptual structure of the method as closely as possible. Rather than a linear procedural workflow, the analysis is represented as a set of dependency-constrained scientific transformations. This provides a direct correspondence between the theoretical stages of MCC and the computational structures used to evaluate them. MCC calculations involve multiple layers of abstraction and heterogeneous computational stages. Entropy contributions arise from different physical observables such as forces, torques, and conformational states, and are evaluated across multiple levels of molecular hierarchy including atoms, residues, and polymers.

Reproducible research software also requires transparent data flow. Intermediate quantities such as covariance matrices, hierarchy assignments, and bead mappings must remain accessible for validation and downstream analysis. These requirements motivate a graph-based architecture in which each scientific transformation is represented explicitly and execution order is derived automatically from data dependencies.

Design principles

The architecture of CodeEntropy is guided by three core design principles:

Method fidelity Major architectural components correspond directly to stages of the MCC method, ensuring that the implementation remains interpretable in terms of the underlying theoretical framework.

Explicit data dependencies. Intermediate quantities used in entropy estimation are represented as explicit objects within the dependency graph, making the flow of scientific data transparent and inspectable.

Scalability through streaming statistics and frame-local execution. The implementation is designed to process long trajectories without storing frame-level observables by accumulating ensemble statistics incrementally, while allowing independent frame-local calculations to be distributed when parallel execution is enabled.

DAG architecture

The CodeEntropy workflow is organised as a composition of DAGs that collectively implement the MCC method. In this model, the computation is expressed as a collection of analysis graphs connected through explicit data dependencies. Nodes represent transformations that derive new physical or statistical quantities from existing ones, while edges represent the required inputs for those transformations. By making dependencies explicit, the DAG structure provides a transparent mapping between MCC theory and its computational implementation.

In practice, the dependency graphs corresponding to static setup, conformational state construction, frame-level statistical accumulation, and entropy estimation are constructed at the beginning of an analysis based on the selected entropy estimators and molecular hierarchy. Together, these graphs define the executable representation of the MCC workflow.

Each node performs a single transformation following the pattern

readcomputewrite.

Nodes retrieve required inputs from the shared data store, perform their transformation, and write results back to the shared state. Each node is placed within the graph according to its dependencies, allowing the scheduler to determine when a node becomes executable. In practice, this corresponds to dependency-resolved execution over the graph, equivalent to evaluation under a topological ordering of valid transformations.

This representation is well suited to MCC because the method itself is naturally defined by transformations whose ordering is constrained by data dependencies. Structural descriptors must exist before frame-level observables can be computed, and entropy estimators depend on statistical quantities accumulated across the trajectory.

In the current implementation, DAGs are used both to make scientific data dependencies explicit and to determine reproducible execution order. The workflow can be executed serially, which remains the simplest reference execution mode, or through an optional Dask-based frame execution layer. In the parallel mode, independent frame-local calculations are distributed across workers after the static setup stage has completed. Each worker evaluates the required frame-level observables for its assigned frames, and the resulting partial statistical contributions are reduced back into the same accumulated ensemble quantities used by the serial workflow.

This design makes the parallelisation boundary explicit: frame-independent topology, hierarchy definitions, bead mappings, and accumulator definitions are constructed once, while frame-dependent observables such as forces, torques, coordinate frames, covariance contributions, and local-environment quantities can be evaluated independently across selected frames. Conformational observations are handled through the specialised conformational pathway described below. The graph-based representation therefore provides both a transparent description of MCC data dependencies and a practical basis for optional parallel frame execution without changing the scientific definition of the calculation.

Mapping MCC theory to DAG nodes

MCC defines a sequence of transformations: construction of a molecular hierarchy, evaluation of frame-dependent observables, accumulation of correlation statistics, and conversion of these statistics into entropy contributions. The CodeEntropy implementation maps each conceptual MCC stage onto one or more dependency-resolved computational nodes.

Static setup nodes construct the frame-independent elements of MCC, including molecular groupings, hierarchy definitions, bead mappings, cached coordinate-frame topology, and persistent statistical accumulators. These objects define the molecular representation and analysis state used during subsequent trajectory processing.

Conformational analysis maps the configurational component of MCC onto dependency-resolved stages for identifying conformational degrees of freedom and estimating conformational state populations.

Frame-processing nodes compute dynamic observables for each selected trajectory frame, including force vectors, torque vectors, frame-level second-moment contributions, and neighbour or local-environment relationships required for orientational entropy estimation. These quantities are incorporated into ensemble statistics through streaming accumulation.

Statistical accumulation nodes update ensemble averages and probability distributions required by MCC estimators, including covariance matrices and conformational state populations.

Finally, estimator nodes transform accumulated statistics into vibrational, configurational, and orientational entropy contributions for each molecule type, hierarchy level, and type of motion.

This alignment between theoretical stages and computational nodes ensures that intermediate quantities correspond directly to well-defined physical observables, simplifying interpretation, validation, and extension of the method.

Execution phases

In a typical analysis, CodeEntropy first constructs the static representation of the molecular system, then performs conformational state construction where required, then processes selected trajectory frames to accumulate statistical observables, and finally evaluates entropy estimators from the accumulated state.

From a software architecture perspective, the workflow is realised as a composition of DAGs with distinct execution semantics rather than as a single monolithic graph. The LevelDAG defines the static setup stage required to initialise molecular hierarchy information, bead mappings, cached topology, and persistent statistical accumulators. A specialised ConformationDAG handles the conformational entropy pathway by organising dihedral topology construction, angle observation, peak detection, and conformational state assignment. The FrameDAG defines the frame-local computations used to evaluate observables and accumulate covariance statistics across selected trajectory frames. Following completion of conformational analysis and frame-level statistical accumulation, a separate EntropyGraph consumes the accumulated state and evaluates the final entropy estimators.

Trajectory access is handled through an explicit frame-selection model. The selected source-frame indices define the frames included in the analysis, and frame-local computation operates on the corresponding frame state. This avoids hidden dependence on trajectory cursor state and makes the relationship between the configured trajectory range and the analysed frames explicit.

The frame-processing stage may be executed serially or using the optional Dask execution layer. In serial execution, selected frames are processed sequentially and incorporated directly into the running statistical state. In optional Dask execution, the selected frame range is divided into independent frame chunks, frame-local observables are evaluated by workers, and partial results are combined through a reduction step. In both execution modes, the same accumulated statistical quantities are produced and subsequently consumed by the EntropyGraph during entropy estimation.

This decomposition separates execution control from scientific data dependencies, exposes the execution semantics of each stage, and provides a clear foundation for reasoning about execution ordering, reproducibility, optimisation, and optional parallelisation while preserving a direct correspondence between the MCC method and its software realisation.

This execution model is illustrated in Figure 1, which summarises the relationship between static setup, conformational state construction, frame-level statistical accumulation, and final entropy evaluation. The figure also highlights the separation between frame-independent setup, conformational analysis, frame-local computation, streaming reduction, and post-processing evaluation over accumulated statistical state.

Figure 1

Overview of the CodeEntropy workflow showing static setup, conformational state construction, frame-level statistical accumulation, and final entropy evaluation.

Static setup (LevelDAG)

The static setup stage of the LevelDAG executes once per run and constructs the frame-independent representation of the molecular system required for MCC analysis. During this phase, dependency-resolved setup operations initialise the molecular hierarchy, analysis degrees of freedom, cached topology, and persistent statistical state. These operations include:

  • detection of molecular groupings,

  • construction of hierarchy levels,

  • generation of bead mappings,

  • construction of cached coordinate-frame topology,

  • initialisation of streaming statistical accumulators.

The resulting structural objects remain fixed throughout trajectory processing and define the coordinate systems and statistical representations used by subsequent stages of the workflow.

Conformational state construction (ConformationDAG)

The ConformationDAG handles the conformational entropy pathway. It transforms molecular topology and trajectory-derived dihedral observations into discrete conformational states. This stage includes construction of dihedral topology, observation of frame-local dihedral angles, peak detection, state assignment, and estimation of conformational state populations. The resulting conformational state information is retained for use by the configurational entropy estimator.

Frame processing (FrameDAG)

The FrameDAG defines the frame-local computations evaluated for each selected trajectory frame, either sequentially in serial execution or across frame chunks in optional Dask execution. For frame f, let Ff and τf denote the force and torque vectors, respectively, and let 𝖳 denote matrix transpose. For each frame, second-moment matrices are constructed as

9
MFf=Ff(Ff)𝖳

and

10
Mτf=τf(τf)𝖳.

These matrices represent instantaneous outer products for a single frame and provide the contributions required to construct ensemble-averaged second moments. They are not retained individually; instead, they are incorporated into running averages across frames using incremental streaming estimators. The resulting ensemble-averaged second moments are then used to compute covariance matrices for entropy estimation.

In addition to force and torque-based observables, the FrameDAG evaluates frame-local neighbour relationships and molecular environments required for orientational entropy estimation. Because these relationships depend on the instantaneous molecular configuration, they are evaluated dynamically during trajectory processing rather than during static setup.

Entropy evaluation (EntropyGraph)

Once conformational state information and accumulated statistical quantities have been produced, the EntropyGraph evaluates the final entropy estimators. It consumes the accumulated covariance statistics, conformational state populations, and other analysis outputs to produce vibrational, orientational, and configurational entropy contributions.

Shared data model

Nodes communicate through a shared mutable state container referred to as shared_data.

Shared data model

This shared data structure acts as a lightweight state container through which nodes exchange intermediate scientific objects without introducing tight coupling between workflow components. The design favours transparency and modularity over heavy abstraction: intermediate scientific objects remain explicitly represented and can therefore be inspected, validated, and serialised during execution.

The shared state also distinguishes between frame-independent and frame-dependent data. Static molecular topology, hierarchy definitions, bead mappings, and cached axes topology are constructed once before frame processing. Conformational topology, conformational observations, and state populations are handled through the specialised conformational pathway. Frame-dependent values such as positions, forces, torques, coordinate axes, covariance contributions, and neighbour relationships are computed during frame processing and accumulated into persistent statistical state. Conformational observations and state populations are handled through the specialised conformational pathway.

Streaming statistical accumulation

MD trajectories may contain very large numbers of frames, often extending to millions of time steps, making storage of per-frame observables impractical. Instead, CodeEntropy accumulates statistics incrementally using streaming estimators. For each frame f, the second moment

11
MFf=Ff(Ff)𝖳

is incorporated into a running mean

12
Mn=Mn1+XnMn1n.

This allows ensemble averages such as FF𝖳 to be computed without storing individual frames. The covariance matrix used in MCC estimators is then obtained as

13
C=FF𝖳FF𝖳.

This streaming formulation allows the DAG-based workflow to operate on long MD trajectories while maintaining bounded memory usage and preserving reproducible statistical accumulation semantics.

Computational characteristics

Let Nf denote the number of selected trajectory frames and k the dimensionality of a force or torque block corresponding to a given set of degrees of freedom. Computing an outer-product second moment for one frame requires O(k2) operations, giving a total cost of O(Nfk2) for streaming accumulation across the trajectory. Eigen decomposition-based entropy estimators scale as O(k3) per block. Because only accumulated statistics are stored, memory usage scales with the dimensionality of the retained matrices rather than the number of frames.

The implementation reduces avoidable per-frame overhead by separating static topology discovery from frame-local numerical work. Molecular hierarchy definitions, bead membership, conformational topology, and coordinate-frame topology are established before frame iteration. The frame loop then reuses these static descriptions while computing only quantities that depend on instantaneous coordinates and forces.

For workflows where frame processing dominates runtime, the optional Dask execution layer can distribute independent frame-local calculations across workers. Partial frame results are reduced into the same accumulated statistical quantities used by the serial workflow, preserving the statistical semantics of the MCC calculation while providing a route to improved performance on multicore or cluster resources.

Output and provenance

Results are written as structured JSON output.

JSON output

This structured output format allows results and associated metadata to be parsed directly by downstream software, supporting reproducible workflows and reuse of derived quantities.

To enhance transparency during execution, results are also rendered interactively in the terminal using the rich library, which produces colour-coded tabular summaries of residue- and molecule-level entropy contributions. Figure 2 shows the command-line interface displayed during program initialisation, while Figure 3 illustrates the YAML-based runtime configuration used to define analysis parameters. Representative terminal outputs showing aggregated entropy summaries and residue-level entropy decomposition are shown in Figures 4 and 5, respectively.

Figure 2

Command-line interface splash screen displayed during program initialisation, showing software identity and execution context.

Figure 3

Example runtime configuration loaded from a YAML file, illustrating user-defined analysis parameters and workflow options.

Figure 4

Terminal output showing aggregated entropy results, including total and component contributions computed from the accumulated statistical model.

Figure 5

Residue-level entropy decomposition presented as tabulated terminal output, illustrating how total entropy contributions are distributed across molecular components.

In addition to structured result export, CodeEntropy records detailed runtime provenance through a configurable logging subsystem. Multiple logging handlers capture complementary aspects of program execution, including interactive console output, command-line invocation, routine informational messages, and error reports. Low-level trajectory processing events originating from MDAnalysis are also recorded to facilitate debugging of data parsing and selection operations.

All log records include timestamps, message severity levels, and source identifiers, enabling reconstruction of the full computational history of an analysis. Together with the structured output files, these logs provide a self-describing record of each run, supporting reproducibility and alignment with FAIR research software principles. The source code is maintained in a public Git repository with versioned releases to PyPI and to the CCPBioSim Conda channel, allowing specific software versions to be associated with published analyses.

Quality control

Quality assurance is integrated throughout the development of CodeEntropy. The node-based architecture enables validation at multiple levels of granularity. Unit tests verify individual nodes using controlled inputs to confirm the correctness of specific analytical stages.

Regression tests are used to ensure numerical stability of entropy calculations across software updates. Representative molecular systems and trajectories are analysed as part of the automated test suite, and resulting entropy components are compared against stored reference values within floating-point tolerance.

Continuous integration workflows automatically execute the test suite whenever changes are introduced to the codebase, ensuring that modifications to the architecture or numerical routines do not introduce unintended deviations in entropy estimates. Together, these practices support the role of CodeEntropy as a reliable, transparent, and reproducible implementation of MCC.

Extensibility

The dependency-graph architecture provides a clear foundation for extending MCC analysis. New entropy estimators can be introduced as additional terminal nodes consuming existing statistical objects, while alternative hierarchy definitions or statistical accumulation strategies can be incorporated by modifying upstream nodes without restructuring the overall workflow. This separation of analytical stages supports both methodological development and long-term software maintainability.

Case studies

We first demonstrate CodeEntropy for a set of 49 liquids which had previously been examined with MCC as detailed elsewhere (4). The trajectories are 1 ns, 1000-frame MD trajectories using the AMBER 14 (23) MD engine and the GAFF force field (24). The calculated entropies in Figure 6 are in reasonable agreement with experiment, with a mean unsigned error (MUE) versus experiment of 9.0 J mol-1 K-1. The mean unsigned difference between MCC values and the previous work (4) is 5.5 J mol-1 K-1. To our knowledge Two-Phase Thermodynamic (2PT) is the only method that has been used to calculate entropy from MD simulations for a range of liquids (25). Similar to what was found earlier (4), for the 12 liquids in common the MUE for 2PT versus experiment is 24.4 J mol-1 K-1, more than triple that of MCC at 7.8 J mol-1 K-1.

Figure 6

Entropy of 49 liquids from CodeEntropy versus experiment. The dashed line represents perfect agreement. Error bars are negligible.

We next demonstrate CodeEntropy for seven host-guest systems from the SAMPL8 (Statistical Assessment of the Modeling of Proteins and Ligands) “Drugs of Abuse” Blind Challenge (26), which had been previously investigated with MCC (16). This used an earlier version of CodeEntropy (12, 27) for the solutes and an earlier version of waterEntropy for water called POSEIDON (15, 28), which had united-atom axes defined using bonded hydrogens instead of bonded united atoms as in the current version, and the origin was at the oxygen for water. MD simulations were run in triplicate with GROMACS 2023.3 (29) using the same starting structures, conditions, and force field, which is GAFF2 (30) for the solutes and TIP3P (31) for water. Entropy was calculated using 1001 frames for the solvated guest, host, and host-guest complexes and only 11 frames for the bulk water systems, which converge much faster with many identical molecules. Terms were included for the positional entropy of the solute in a 1 M solution and orientational entropy of the solute when bound, which are discussed elsewhere in more detail (16) but not yet included in CodeEntropy. For water, the factor of 1/4 in Ωmorient for hydrogen-bond correlation (3) is not needed because it would entirely cancel in the entropy difference. The free energy of binding is calculated using the standard relationship ΔG=ΔHTΔS where the differences in enthalpy ΔH and entropy ΔS are calculated from the values for the bound complex and bulk water minus those of the unbound host and guest. Figure 7 shows satisfactory agreement with experiment for ΔG with a MUE versus experiment of 2.4 kcal mol-1 and a MUE relative to the previous work (16) of 2.2 kcal mol-1. Errors are inevitably larger, given the reduced symmetry compared to pure liquids and the large number of degrees of freedom, especially from the solvent water. For guest G2, the larger error in ΔG is partly from the force-field energy and partly from insufficiently sampled binding poses of the guest in the host, issues which are beyond the scope of this work but highlight the need for multiple simulations. The MUE here versus experiment is not as close as the MUE in previous work (16), which was 0.9 kcal mol-1, but is comparable to the error bars and demonstrates that longer simulations and other techniques will be needed to reduce errors. For the SAMPL8 challenge,(26) the MCC results are close to the three MM/PBSA entries (Molecular Mechanics/Poisson Boltzmann Surface Area), a method that also predicts free energy directly, which had MUEs of 1.7 kcal mol-1 but fell short of the top-ranked alchemical method using Bennett’s acceptance ratio and a polarisable force field which had a MUE <1 kcal mol-1. Incidentally, Ali et al. (4) achieved the top-ranked ΔH prediction in SAMPL8, but that does not involve MCC.

Figure 7

CodeEntropy binding Gibbs energies ΔG (blue) versus experiment (green) for the host-guest systems in the SAMPL8 challenge.

CodeEntropy is applicable to polymers such as proteins and DNA, but examples are not included here, owing to a lack of experimental entropies for comparison. Further refinements to the configurational entropy, levels of hierarchy and overall accuracy are in progress.

(2) Availability

Operating System

The software has been tested on the following operating systems:

  • Ubuntu 24.04 LTS

  • macOS Sequoia

  • Windows 11

Programming Language

This codebase has been written in Python 3.12 or newer. Older versions of Python will not work due to dependency and syntax requirements.

Additional System Requirements

No specialised hardware is required. The software runs on standard desktop or workstation hardware. Memory usage depends on the size of the molecular system and trajectory length, as statistical accumulators scale with the dimensionality of the molecular representation rather than the number of frames.

Dependencies

The software depends on the following Python packages:

  • numpy (2.3, <3.0)

  • MDAnalysis (2.10, <3.0)

  • pandas (3.0, <3.1)

  • psutil (7.1, <7.0)

  • PyYAML (6.0, <7.0)

  • python-json-logger 4.0, <5.0

  • rich (14.2, <15)

  • art (6.5 <7.0)

  • networkx (3.6, <3.7)

  • matplotlib (3.11, <3.12)

  • waterEntropy (2, <2.3)

  • requests (2.32, <3.0)

  • rdkit (2025.9.5)

  • numba 0.65, <0.70

List of Contributors

  • Arghya Chakravorty

  • Donald Chung

  • Sarah Fegan

  • James Gebbie-Rayet

  • Sarah Harris

  • Richard Henchman

  • Jonathan Higham

  • Jas Kalayan

  • Ioana Papa

  • Harry Swift

Software location

Archive

Code repository

Language

English.

(3) Reuse Potential

CodeEntropy enables reproducible and extensible entropy analysis across a broad range of molecular systems. It provides a unified framework for applying the MCC method to MD simulations, supporting studies of liquids, biomolecules, and molecular complexes. Researchers in computational chemistry, biophysics, and materials science can reuse the software to quantify entropy contributions to molecular stability, ligand binding, and solvation processes, or to benchmark alternative theoretical approaches for entropy estimation.

Beyond direct application, CodeEntropy can serve as a reference implementation for methodological development. Its transparent architecture and open-source availability make it suitable for educational purposes, code comparison studies, or integration into larger simulation and analysis workflows. Because the codebase emphasises clarity, modularity, and configuration-driven workflows, it also provides a foundation for implementing new entropy-based models or coupling MCC calculations with related approaches for free-energy estimation.

The software is openly available under the MIT License through a public GitHub repository, where users can access documentation, example datasets, and automated test suites. The repository supports collaborative development through issue tracking and pull requests, allowing the software to evolve as a community resource for reliable and reproducible entropy estimation in molecular simulation.

Acknowledgements

We would like to thank Hafiz Saqib Ali for sharing his MD trajectory data for the liquid and host-guest systems.

DOI: https://doi.org/10.5334/jors.741 | Journal eISSN: 2049-9647
Language: English
Page range: 55 - 55
Submitted on: May 13, 2026
Accepted on: Jul 7, 2026
Published on: Jul 20, 2026
Published by: Ubiquity Press
In partnership with: Paradigm Publishing Services

© 2026 Harry Swift, Jas Kalayan, Ioana A. Papa, Sarah K. Fegan, Richard H. Henchman, Sarah A. Harris, James Gebbie-Rayet, published by Ubiquity Press
This work is licensed under the Creative Commons Attribution 4.0 License.