pheval
Details
| GitHub | monarch-initiative/pheval |
| Language | Python |
| Description | A framework for empirical evaluation of phenotype matching and prioritisation |
Dependencies
External Dependencies
| Package | Version |
|---|---|
| jaydebeapi | >=1.2.3 |
| tqdm | >=4.64.1 |
| pandas | >=1.5.1 |
| deprecation | >=2.1.0 |
| click | >=8.1.3 |
| class-resolver | >=0.4.2 |
| phenopackets | >=2.0.2,<3 |
| oaklib | >=0.5.6 |
| >=3.0.0,<4 | |
| pyaml | >=21.10.1,<22 |
| plotly | >=5.13.0,<6 |
| seaborn | >=0.12.2,<0.13 |
| matplotlib | >=3.7.0,<4 |
| pyserde | >=0.9.8,<0.10 |
| polars | ~=1.23 |
| scikit-learn | >=1.4.0,<2 |
| duckdb | >=1.0.0,<2 |
| pyarrow | >=19.0.1,<20 |
Documentation
PhEval - Phenotypic Inference Evaluation Framework
PhEval (Phenotypic Inference Evaluation Framework) is a modular, reproducible benchmarking framework for evaluating phenotype-driven prioritisation tools, such as gene, variant, and disease prioritisation algorithms.
It is designed to support fair comparison across tools, tool versions, datasets, and knowledge updates, addressing a long-standing gap in standardised evaluation for phenotype-based methods.
📖 Full documentation: https://monarch-initiative.github.io/pheval/
Why PhEval?
Evaluating phenotype-driven prioritisation tools is challenging because performance depends on many moving parts, including:
- Phenotype representations and noise
- Ontology structure and versioning
- Gene and disease mappings
- Tool-specific scoring and ranking strategies
- Input cohorts and simulation approaches
PhEval provides a framework that makes these factors explicit, controlled, and comparable.
Key features:
- Standardised outputs across tools
- Reproducible benchmarking with recorded metadata
- Plugin-based architecture for extensibility
- Separation of execution and evaluation
- Support for gene, variant, and disease prioritisation
Installation
PhEval requires Python 3.10 or later.
Install from PyPI:
pip install pheval
This installs:
- The core pheval CLI (for running tools via plugins)
pheval-utils(for data preparation, benchmarking, and analysis)
Verify installation:
pheval --help
pheval-utils --help
How PhEval is used
PhEval workflows typically consist of three phases:
- Prepare data Prepare and manipulate phenopackets and related inputs (e.g. VCFs).
- Run tools
Execute phenotype-driven prioritisation tools via plugin-provided runners using:
pheval run --runner <runner_name> ... - Benchmark and analyse Compare results across runs using standardised metrics and plots.
Each phase is documented in detail in the user documentation.
Plugins and runners
PhEval itself is tool-agnostic.
Support for specific tools is provided via plugins, which implement runners responsible for:
- Preparing tool inputs
- Executing the tool
- Converting raw outputs into PhEval standardised results
A list of available plugins is maintained in the documentation:
Plugins: https://monarch-initiative.github.io/pheval/plugins/
Each plugin repository contains tool-specific installation instructions and examples.
Documentation
The PhEval documentation is organised by audience and task: * Getting started: installation and first steps * Using PhEval: running tools, plugins, and workflows * Utilities: data preparation, phenopacket manipulation, simulations * Benchmarking: executing benchmarks, metrics, and plots * Developer documentation: plugin development and API reference
Start here: https://monarch-initiative.github.io/pheval/
Contributions
Contributions are welcome across:
- Code
- Documentation
- Testing
- Plugins and integrations
Citation
If you use PhEval in your research, please cite the following publication:
Bridges, Y., Souza, V. d., Cortes, K. G., et al.
Towards a standard benchmark for phenotype-driven variant and gene prioritisation algorithms: PhEval – Phenotypic Inference Evaluation Framework.
BMC Bioinformatics 26, 87 (2025).
https://doi.org/10.1186/s12859-025-06105-4