Skip to content

Valentine

(Schema-) Matching DataFrames Made Easy.

Get started API reference View on GitHub

  • Coma


    Schema + instances. General-purpose first choice — strong defaults, informative sub-scores.

  • Cupid


    Schema only. Nested schemas where column names and structure matter more than data.

  • DistributionBased


    Instances only. Matching by value distributions when names are unreliable.

  • JaccardDistanceMatcher


    Instances only. Simple, explainable baseline — useful for sanity checks.

  • SimilarityFlooding


    Schema only. Structure-heavy schemas where graph neighbourhoods carry signal.

PyPI version Python versions PyPI downloads Build codecov License

Valentine is a Python package for capturing potential relationships among columns of different tabular datasets, given as pandas or Polars DataFrames. It implements several schema- and instance-based matching algorithms behind a single, uniform API, and ships with evaluation metrics so you can measure match quality against a ground truth. Pandas and Polars frames can be freely mixed in the same call.

Installation

pip install valentine             # pandas only
pip install valentine[polars]     # pandas + Polars support
pip install valentine[embeddings] # pandas + embeddings support

Requires Python >=3.10, <3.15.

A 30-second taste

import pandas as pd
from valentine import valentine_match
from valentine.algorithms import Coma

df1 = pd.read_csv("source_candidates.csv")
df2 = pd.read_csv("target_candidates.csv")

matches = valentine_match([df1, df2], Coma(use_instances=True))

for pair, score in matches.items():
    print(f"{pair.source_column} <-> {pair.target_column}: {score:.3f}")

Ready for more? Head over to Getting started, or jump straight to the API reference.

Research

Valentine started as a research project at TU Delft and is based on the ICDE 2021 paper. See the Research page for the papers behind the package, the algorithms it implements, and citation info.