SARVault is an exploration of cheminformatics at warehouse scale. The guiding questions are the ones a medicinal chemistry team working on cytotoxic ADC payloads keeps asking: how does potency rank across the series, where are the activity cliffs, which analogues exist, what does the chemical neighbourhood look like?
The raw material is public - ChEMBL, UniChem, PDBe - but fragmented across sources with different identifiers, formats, and update cycles. Integrating heterogeneous sources into one consistent, queryable model is the same problem that dominates any enterprise informatics deployment. SARVault works this problem end to end. Every step of the pipeline is reproducible.
Scale is not what makes this hard. Under two thousand compounds fit in memory on any laptop. What makes it hard is that the same molecule wears a different name in every source, while a potency value means nothing without the assay it came from. Both problems stay invisible until an answer is already wrong. A join that silently misses is worse than a join that fails.