Skip to content

Latest commit

 

History

26 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Compresso: A PyTorch Framework for Sparse Representation Learning

PyPI Python License Docs Live Demo Documentation build

Compresso logo

Compresso is an open-source PyTorch framework for sparse representation learning. It provides reusable building blocks for learning sparse neural representations, dynamic sparsification, sparse inference, and semantic analysis, enabling researchers to rapidly prototype sparse neural architectures while focusing on models rather than infrastructure.

Why Compresso?

Sparse representations are becoming increasingly important across machine learning due to their efficiency, interpretability, and ability to capture semantically meaningful concepts. Yet building sparse models often requires implementing pruning schedules, sparse kernels, training loops, device management, and visualization from scratch.

Compresso hides this complexity behind a simple, modular API.

The name is inspired by Italian espresso culture: when you order a coffee in Italy, you simply ask for a caffè. The barista handles the beans, pressure, and brewing; you just enjoy the result. Compresso follows the same philosophy: researchers should be able to train and analyze sparse representations without worrying about the underlying engineering.

Install

Using pip:

pip install compresso-pytorch

For local development:

git clone https://github.com/zombak79/compresso.git
cd compresso
pip install -e ".[test]"

Documentation

Documentation is available at https://zombak79.github.io/compresso/.

Minimal Example

You can train a sparse autoencoder through one high-level class TopKSAETrainer with a scikit-learn-style wrapper: fit trains on a dense matrix, transform returns sparse codes, and fit_transform does both. All hyperparameters live in the TopKSAEConfig dataclass.

import numpy as np
from compresso import TopKSAEConfig, TopKSAETrainer

embeddings = np.random.randn(10_000, 512).astype("float32")

trainer = TopKSAETrainer(
    TopKSAEConfig(
        hidden_dim=4096,
        k=32,
    )
)

srp = trainer.fit_transform(embeddings)
print(srp)

To train a denoising SAE, enable Gaussian corruption. The trainer adds noise only to training inputs and still reconstructs the original clean embeddings:

trainer = TopKSAETrainer(
    TopKSAEConfig(
        hidden_dim=4096,
        k=32,
        noise_type="gaussian",
    )
)
srp = trainer.fit_transform(embeddings)

Clustering and cluster labeling can be run through the clustering pipeline:

from compresso import clustering as cc

cluster_graph = cc.ClusteringPipeline(
    [
        cc.DominantSignedClustering(min_cluster_size=20),
        cc.LabelClusters(...),
    ]
)(srp)

See full example at https://zombak79.github.io/compresso/clustering.html.

Recommender Systems Add-on

For recommender-system experiments, see compresso-recsys, the companion package built on top of Compresso.

It provides recommender-specific dataset loaders, checkpoint management, and retrieval metrics such as Recall and nDCG. It can be installed with:

pip install compresso-recsys

Compresso contains the general sparse representation learning components, while compresso-recsys provides the infrastructure needed to apply and evaluate them in recommender-system experiments.

Citation

If you find this project helpful or use it in your academic work, please consider citing it. This helps us continue to maintain and develop this project. You can find the citation format below.

For method-specific references, including the sparse embedding compression work behind TopKSAETrainer, see the citation guide.

@inproceedings{10.1145/3773078.3841254,
  author = {Van{\v c}ura, Vojt{\v e}ch and Medda, Giacomo and Spi{\v s}{\'a}k, Martin and Pe{\v s}ka, Ladislav},
  title = {COMPRESSO: Espresso-Style Sparse Representation Learning for Interpretable Recommender Systems},
  year = {2026},
  isbn = {9798400722844},
  publisher = {Association for Computing Machinery},
  address = {New York, NY, USA},
  url = {https://doi.org/10.1145/3773078.3841254},
  doi = {10.1145/3773078.3841254},
  abstract = {Sparse representations can make recommender-system embeddings more compact and inspectable, but developing sparse-learning workflows typically requires substantial engineering around sparsification, training, pruning, storage, and analysis. We present Compresso, an open-source PyTorch framework that exposes this functionality through a simple and modular interface. Inspired by Italian espresso culture, where one orders a caff{\`e} while the barista handles the machinery, Compresso lets researchers focus on sparse models rather than infrastructure. The framework provides high-level sparse autoencoder training, reusable sparse tensor representations, differentiable top-k operators, sparse and masked neural parameters, pruning schedules, and composable clustering tools.},
  booktitle = {Proceedings of the 20th ACM Conference on Recommender Systems},
  pages = {1793–1795},
  numpages = {3},
  keywords = {Sparse Representations, Embedding Compression, Interpretability},
  location = {
  },
  series = {RecSys '26}
}