A tensor-network extension for Hamiltonians from the Azulene Labs / HamLib workflow, with validated CPU execution and optional GPU acceleration.
hamlib_functions provides helper functions for the
HamLib Hamiltonian library.
The optional hamlib.tensor_networks package continues the existing data path:
HamLib Hamiltonian
→ hamlib.hamlib_snippets
→ OpenFermion QubitOperator / explicit mat2qubit conversion
→ canonical Pauli terms
→ matrix-product operator (MPO)
→ matrix-product state (MPS)
→ observables / two-site DMRG
→ NumPy CPU or optional CuPy CUDA execution
It does not introduce another Hamiltonian format. HamLib HDF5 data is loaded by
the existing read_openfermion_hdf5 or read_mat2qubit_hdf5 functions.
Matrix and d-level encodings continue to use mat2qubit, whose output is an
OpenFermion QubitOperator. Fermionic transforms and mat2qubit subsystem
metadata are explicit rather than guessed.
- HamLib-compatible
QubitOperator,FermionOperator, andqSymbOpadapters. - Canonical, complex-valued Pauli-term representation.
- Pauli-string to MPO conversion and QR/SVD compression with tolerance control.
- Product and deterministic random MPS initialization, application, and compression.
- Expectation values, energy variance, and finite two-site DMRG.
- NumPy/SciPy reference execution and optional CuPy CUDA observables.
- Exact dense validation for small systems.
- Reproducible CSV/JSON benchmarks, plots, and regression comparison.
There is currently no Triton kernel, OpenMP extension, or GPU DMRG eigensolver.
This repository currently has no pyproject.toml or setup.py; it is used from
a source checkout. The commands below reflect that existing layout.
git clone https://github.com/Azulene-Labs/hamlib_functions.git
git clone https://github.com/Azulene-Labs/mat2qubit.git
cd hamlib_functions
python -m venv .venv
. .venv/bin/activate
python -m pip install numpy scipy h5py networkx openfermion qiskit pytest
python -m pip install -e ../mat2qubit
export PYTHONPATH="$PWD${PYTHONPATH:+:$PYTHONPATH}"
python -m pytest -qInstall a CuPy package or wheel matching the host CUDA toolkit and NVIDIA driver
using the CuPy installation selector.
A source installation can be requested with python -m pip install cupy, though
a compatible prebuilt wheel is normally preferable. This project deliberately
does not hardcode a CUDA wheel suffix.
nvidia-smi
python -c "import cupy; print(cupy.cuda.runtime.getDeviceCount())"The default installation remains CPU-only. Triton and OpenMP installation sections are omitted because those backends are not implemented.
from hamlib import hamlib_snippets
from hamlib.tensor_networks import (
energy_variance_reference,
pauli_terms_to_mpo,
product_state_mps,
qubit_operator_to_pauli_terms,
)
# Load a Pauli Hamiltonian through the existing HamLib API.
key = "graph-1D-grid-nonpbc-qubitnodes_Lx-4_h-1"
hamiltonian = hamlib_snippets.read_openfermion_hdf5("tfim.hdf5", key)
terms = qubit_operator_to_pauli_terms(hamiltonian)
mpo = pauli_terms_to_mpo(terms, n_qubits=4, compression_tolerance=1e-12)
# Evaluate the energy and variance of |0000>.
mps = product_state_mps([0, 0, 0, 0], dtype="complex128")
energy, variance = energy_variance_reference(mpo, mps)
print(energy, variance)The complete official-TFIM and DMRG example is
examples/tfim_tensor_network.py.
Transfers are explicit and occur once before both observables:
import cupy as cp
from hamlib.tensor_networks import energy_variance, pauli_terms_to_mpo, random_mps
if cp.cuda.runtime.getDeviceCount() < 1:
raise RuntimeError("CUDA is unavailable")
device, dtype = "cuda:0", "complex128"
terms = [(1.0, [(0, "Z"), (1, "Z")]), (0.5, [(0, "X")])]
cpu_mpo = pauli_terms_to_mpo(terms, n_qubits=2, dtype=dtype)
cpu_mps = random_mps(2, bond_dimension=2, dtype=dtype, seed=7)
cpu_energy, cpu_variance = energy_variance(cpu_mpo, cpu_mps)
with cp.cuda.Device(0):
gpu_mpo = [cp.asarray(tensor) for tensor in cpu_mpo]
gpu_mps = [cp.asarray(tensor) for tensor in cpu_mps]
gpu_energy, gpu_variance = energy_variance(
gpu_mpo, gpu_mps, "cupy", device, dtype
)
errors = cp.asnumpy(cp.abs(cp.asarray([
gpu_energy - cpu_energy, gpu_variance - cpu_variance
])))
print({"energy": gpu_energy, "variance": gpu_variance,
"max_error": errors.max(), "backend": "cupy",
"device": device, "dtype": dtype})This API does not imply that a GPU is faster. Transfer overhead and many small site contractions can dominate small problems. On the checked-in RTX 4060 Laptop GPU run, both the CuPy kernel-only and end-to-end paths were slower than NumPy for the tested low-bond TFIM sizes.
Canonical terms use (coefficient, [(qubit_index, pauli_character), ...]).
Indices are zero-based. Missing qubits imply identity; identity has an empty
factor list. Factors and terms are sorted, duplicates are combined, and complex
coefficients are preserved. Numerical filtering happens after combination.
An MPO tensor has shape (left_bond, right_bond, output, input) and an MPS
tensor has shape (left_bond, right_bond, physical). Open boundaries require
the first left bond and final right bond to equal one. MPO application contracts
[ B_{(a,l),(b,r)}^{s} = \sum_t W_{l,r}^{s,t} A_{a,b}^{t}. ]
Dense validation orders sites as (|q_0 q_1 \ldots q_{n-1}\rangle).
mpo_to_dense(..., reverse_qubits=True) matches mat2qubit's default convention,
where qubit zero is the rightmost Kronecker factor.
compression_tolerance is an absolute local singular-value cutoff. None
skips rounding and retains the product-sum MPO; 0.0 rounds without
intentionally dropping nonzero singular values. DMRG reports accumulated
discarded singular-value weight as truncation_error; neither value is a
universal bound on energy error.
Small systems are reconstructed densely. MPO application is compared with
dense matrix-vector application, and DMRG energy with exact diagonalization.
Default tolerances are atol=1e-10, rtol=1e-8 for float64/complex128 and
atol=1e-5, rtol=1e-4 for float32/complex64.
Measured by run 20260823T122039Z:
| Test | System | Reference | Max error | Status |
|---|---|---|---|---|
| MPO reconstruction | 6 qubits | dense matrix | 0 | PASS |
| MPO–MPS application | 6 qubits | dense vector | 5.55e-16 | PASS |
| DMRG ground energy | 8 qubits | exact diagonalization | 3.73e-14 | PASS |
| GPU observables | 4–12 qubits | CPU reference | 7.76e-15 | PASS |
See results/tables/benchmark_summary.md
for the generated table.
The checked-in run used official NERSC TFIM entries for 4–12 qubits, a 13th
Gen Intel Core i7-13700H under WSL2, Python 3.10.12, NumPy 2.2.6, SciPy 1.15.3,
OpenFermion 1.8.1, complex128 tensors, MPS bond dimension 16, one configured
native thread, three warm-ups, and ten timed iterations. CPU timing used
time.perf_counter.
The GPU was an NVIDIA GeForce RTX 4060 Laptop GPU with 8 GB VRAM, NVIDIA driver 596.08, CuPy 14.1.1, and CUDA runtime 12.9. Synchronized CUDA events timed kernels; synchronized wall time measured explicit transfers. At 12 qubits, MPO–MPS application took 0.624 ms on CPU, 2.249 ms for the CUDA kernel path, and 3.889 ms end-to-end. These correspond to speedups of 0.278× and 0.161×—that is, CUDA was slower for this small, low-bond workload.
A separate 20-qubit crossover study increased the MPS bond dimension. CUDA kernel speedup was 0.779× at bond 32, 1.237× at bond 48, and 2.419× at bond 64. End-to-end speedup remained below one because each measurement included explicit host/device transfers. Reusing GPU-resident tensors is therefore necessary to benefit from the measured kernel crossover.
A system-size sweep at fixed bond dimension 64 covered 10, 12, 14, 16, 18, and
20 qubits. The CUDA kernel crossed over at 16 qubits (1.260×), reached 1.800× at
18 qubits, and measured 1.740× at 20 qubits. End-to-end speedup remained below
one at every size; its best measured value was 0.803× at 18 qubits. These are
official HamLib TFIM cases, not synthetic matrix benchmarks. See the generated
bond-64 summary.
The extended 4–24-qubit MPO-approximation run used compression tolerance
1e-10, which held the TFIM MPO bond dimension at 3. At 24 qubits, the CUDA
kernel was 2.642× faster and end-to-end execution was 1.002× faster; the latter
is effectively break-even at this timing precision. Output allocation was
15.0 MiB. The small dense validation measured 9.77e-15 maximum operator error
at this tolerance. See the generated
4–24-qubit summary.
python benchmarks/run_benchmarks.py --hamlib-hdf5 /path/to/tfim.hdf5
python benchmarks/plot_results.py
python benchmarks/summarize_results.pySee benchmarks/README.md for methodology, CUDA timing,
regression checks, and Nsight commands. See
docs/performance_report.md for measured results
and GPU-programming priorities.
- MPS efficiency depends on entanglement and available bond dimension.
- Orbital and qubit ordering can materially change bond dimensions.
- Small contractions may be faster on CPU; GPU transfer overhead can dominate.
- Complex128 GPU throughput can differ substantially from complex64 throughput.
- Triton may not outperform CuPy/vendor libraries for general contractions.
- Exact dense validation is exponentially limited to small systems.
- DMRG is a clarity-first research/reference implementation using SciPy on CPU.
- Process peak memory is not reported without an allocator-aware monitor; generated tensor allocation is reported instead.
- Native HamLib/OpenFermion adapter
- Explicit mat2qubit and fermionic conversion path
- CPU MPO construction and QR/SVD compression
- CPU MPS observables and variance
- Two-site DMRG solver
- Dense CPU validation
- Optional CuPy CUDA MPO application and observables
- Reproducible CSV/JSON benchmarks and plots
- Validate and benchmark CUDA on an RTX 4060 Laptop GPU
- GPU-resident DMRG effective eigensolver
- Profile-driven grouped/batched GPU contractions
- Profile-driven Triton kernel
- Optional OpenMP backend
- Multi-GPU execution
- Molecular orbital-ordering heuristics
HamLib is described by Sawaya et al. in arXiv:2306.13126. Preserve the citations and licenses of HamLib, OpenFermion, and mat2qubit when using their data or methods.
hamlib_functions is maintained by Azulene Labs and NERSC contributors Nicolas
Sawaya, Katie Klymko, and Daan Camps. This tensor-network contribution extends
their public APIs; it does not imply additional affiliation or endorsement.













