Skip to content

Latest commit

 

History

12 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

HamLib functions

A tensor-network extension for Hamiltonians from the Azulene Labs / HamLib workflow, with validated CPU execution and optional GPU acceleration.

Overview

hamlib_functions provides helper functions for the HamLib Hamiltonian library. The optional hamlib.tensor_networks package continues the existing data path:

HamLib Hamiltonian
  → hamlib.hamlib_snippets
  → OpenFermion QubitOperator / explicit mat2qubit conversion
  → canonical Pauli terms
  → matrix-product operator (MPO)
  → matrix-product state (MPS)
  → observables / two-site DMRG
  → NumPy CPU or optional CuPy CUDA execution

It does not introduce another Hamiltonian format. HamLib HDF5 data is loaded by the existing read_openfermion_hdf5 or read_mat2qubit_hdf5 functions. Matrix and d-level encodings continue to use mat2qubit, whose output is an OpenFermion QubitOperator. Fermionic transforms and mat2qubit subsystem metadata are explicit rather than guessed.

Features

  • HamLib-compatible QubitOperator, FermionOperator, and qSymbOp adapters.
  • Canonical, complex-valued Pauli-term representation.
  • Pauli-string to MPO conversion and QR/SVD compression with tolerance control.
  • Product and deterministic random MPS initialization, application, and compression.
  • Expectation values, energy variance, and finite two-site DMRG.
  • NumPy/SciPy reference execution and optional CuPy CUDA observables.
  • Exact dense validation for small systems.
  • Reproducible CSV/JSON benchmarks, plots, and regression comparison.

There is currently no Triton kernel, OpenMP extension, or GPU DMRG eigensolver.

Installation

This repository currently has no pyproject.toml or setup.py; it is used from a source checkout. The commands below reflect that existing layout.

CPU-only installation

git clone https://github.com/Azulene-Labs/hamlib_functions.git
git clone https://github.com/Azulene-Labs/mat2qubit.git
cd hamlib_functions
python -m venv .venv
. .venv/bin/activate
python -m pip install numpy scipy h5py networkx openfermion qiskit pytest
python -m pip install -e ../mat2qubit
export PYTHONPATH="$PWD${PYTHONPATH:+:$PYTHONPATH}"
python -m pytest -q

Optional CUDA installation

Install a CuPy package or wheel matching the host CUDA toolkit and NVIDIA driver using the CuPy installation selector. A source installation can be requested with python -m pip install cupy, though a compatible prebuilt wheel is normally preferable. This project deliberately does not hardcode a CUDA wheel suffix.

nvidia-smi
python -c "import cupy; print(cupy.cuda.runtime.getDeviceCount())"

The default installation remains CPU-only. Triton and OpenMP installation sections are omitted because those backends are not implemented.

Quick start

from hamlib import hamlib_snippets
from hamlib.tensor_networks import (
    energy_variance_reference,
    pauli_terms_to_mpo,
    product_state_mps,
    qubit_operator_to_pauli_terms,
)

# Load a Pauli Hamiltonian through the existing HamLib API.
key = "graph-1D-grid-nonpbc-qubitnodes_Lx-4_h-1"
hamiltonian = hamlib_snippets.read_openfermion_hdf5("tfim.hdf5", key)
terms = qubit_operator_to_pauli_terms(hamiltonian)
mpo = pauli_terms_to_mpo(terms, n_qubits=4, compression_tolerance=1e-12)

# Evaluate the energy and variance of |0000>.
mps = product_state_mps([0, 0, 0, 0], dtype="complex128")
energy, variance = energy_variance_reference(mpo, mps)
print(energy, variance)

The complete official-TFIM and DMRG example is examples/tfim_tensor_network.py.

GPU quick start

Transfers are explicit and occur once before both observables:

import cupy as cp
from hamlib.tensor_networks import energy_variance, pauli_terms_to_mpo, random_mps

if cp.cuda.runtime.getDeviceCount() < 1:
    raise RuntimeError("CUDA is unavailable")

device, dtype = "cuda:0", "complex128"
terms = [(1.0, [(0, "Z"), (1, "Z")]), (0.5, [(0, "X")])]
cpu_mpo = pauli_terms_to_mpo(terms, n_qubits=2, dtype=dtype)
cpu_mps = random_mps(2, bond_dimension=2, dtype=dtype, seed=7)
cpu_energy, cpu_variance = energy_variance(cpu_mpo, cpu_mps)
with cp.cuda.Device(0):
    gpu_mpo = [cp.asarray(tensor) for tensor in cpu_mpo]
    gpu_mps = [cp.asarray(tensor) for tensor in cpu_mps]
    gpu_energy, gpu_variance = energy_variance(
        gpu_mpo, gpu_mps, "cupy", device, dtype
    )
errors = cp.asnumpy(cp.abs(cp.asarray([
    gpu_energy - cpu_energy, gpu_variance - cpu_variance
])))
print({"energy": gpu_energy, "variance": gpu_variance,
       "max_error": errors.max(), "backend": "cupy",
       "device": device, "dtype": dtype})

This API does not imply that a GPU is faster. Transfer overhead and many small site contractions can dominate small problems. On the checked-in RTX 4060 Laptop GPU run, both the CuPy kernel-only and end-to-end paths were slower than NumPy for the tested low-bond TFIM sizes.

Mathematical conventions

Canonical terms use (coefficient, [(qubit_index, pauli_character), ...]). Indices are zero-based. Missing qubits imply identity; identity has an empty factor list. Factors and terms are sorted, duplicates are combined, and complex coefficients are preserved. Numerical filtering happens after combination.

An MPO tensor has shape (left_bond, right_bond, output, input) and an MPS tensor has shape (left_bond, right_bond, physical). Open boundaries require the first left bond and final right bond to equal one. MPO application contracts

[ B_{(a,l),(b,r)}^{s} = \sum_t W_{l,r}^{s,t} A_{a,b}^{t}. ]

Dense validation orders sites as (|q_0 q_1 \ldots q_{n-1}\rangle). mpo_to_dense(..., reverse_qubits=True) matches mat2qubit's default convention, where qubit zero is the rightmost Kronecker factor.

compression_tolerance is an absolute local singular-value cutoff. None skips rounding and retains the product-sum MPO; 0.0 rounds without intentionally dropping nonzero singular values. DMRG reports accumulated discarded singular-value weight as truncation_error; neither value is a universal bound on energy error.

Validation

Small systems are reconstructed densely. MPO application is compared with dense matrix-vector application, and DMRG energy with exact diagonalization. Default tolerances are atol=1e-10, rtol=1e-8 for float64/complex128 and atol=1e-5, rtol=1e-4 for float32/complex64.

Measured by run 20260823T122039Z:

Test System Reference Max error Status
MPO reconstruction 6 qubits dense matrix 0 PASS
MPO–MPS application 6 qubits dense vector 5.55e-16 PASS
DMRG ground energy 8 qubits exact diagonalization 3.73e-14 PASS
GPU observables 4–12 qubits CPU reference 7.76e-15 PASS

See results/tables/benchmark_summary.md for the generated table.

Performance

The checked-in run used official NERSC TFIM entries for 4–12 qubits, a 13th Gen Intel Core i7-13700H under WSL2, Python 3.10.12, NumPy 2.2.6, SciPy 1.15.3, OpenFermion 1.8.1, complex128 tensors, MPS bond dimension 16, one configured native thread, three warm-ups, and ten timed iterations. CPU timing used time.perf_counter.

The GPU was an NVIDIA GeForce RTX 4060 Laptop GPU with 8 GB VRAM, NVIDIA driver 596.08, CuPy 14.1.1, and CUDA runtime 12.9. Synchronized CUDA events timed kernels; synchronized wall time measured explicit transfers. At 12 qubits, MPO–MPS application took 0.624 ms on CPU, 2.249 ms for the CUDA kernel path, and 3.889 ms end-to-end. These correspond to speedups of 0.278× and 0.161×—that is, CUDA was slower for this small, low-bond workload.

A separate 20-qubit crossover study increased the MPS bond dimension. CUDA kernel speedup was 0.779× at bond 32, 1.237× at bond 48, and 2.419× at bond 64. End-to-end speedup remained below one because each measurement included explicit host/device transfers. Reusing GPU-resident tensors is therefore necessary to benefit from the measured kernel crossover.

A system-size sweep at fixed bond dimension 64 covered 10, 12, 14, 16, 18, and 20 qubits. The CUDA kernel crossed over at 16 qubits (1.260×), reached 1.800× at 18 qubits, and measured 1.740× at 20 qubits. End-to-end speedup remained below one at every size; its best measured value was 0.803× at 18 qubits. These are official HamLib TFIM cases, not synthetic matrix benchmarks. See the generated bond-64 summary.

Bond-64 runtime through 20 qubits

Bond-64 CUDA speedup through 20 qubits

The extended 4–24-qubit MPO-approximation run used compression tolerance 1e-10, which held the TFIM MPO bond dimension at 3. At 24 qubits, the CUDA kernel was 2.642× faster and end-to-end execution was 1.002× faster; the latter is effectively break-even at this timing precision. Output allocation was 15.0 MiB. The small dense validation measured 9.77e-15 maximum operator error at this tolerance. See the generated 4–24-qubit summary.

Runtime through 24 qubits

CUDA speedup through 24 qubits

Runtime breakdown through 24 qubits

MPO bond dimension through 24 qubits

Compression accuracy

Output memory through 24 qubits

Runtime versus system size

CPU versus GPU speedup

Runtime breakdown

MPO bond dimension

Accuracy versus compression tolerance

Memory usage versus system size

python benchmarks/run_benchmarks.py --hamlib-hdf5 /path/to/tfim.hdf5
python benchmarks/plot_results.py
python benchmarks/summarize_results.py

See benchmarks/README.md for methodology, CUDA timing, regression checks, and Nsight commands. See docs/performance_report.md for measured results and GPU-programming priorities.

Limitations

  • MPS efficiency depends on entanglement and available bond dimension.
  • Orbital and qubit ordering can materially change bond dimensions.
  • Small contractions may be faster on CPU; GPU transfer overhead can dominate.
  • Complex128 GPU throughput can differ substantially from complex64 throughput.
  • Triton may not outperform CuPy/vendor libraries for general contractions.
  • Exact dense validation is exponentially limited to small systems.
  • DMRG is a clarity-first research/reference implementation using SciPy on CPU.
  • Process peak memory is not reported without an allocator-aware monitor; generated tensor allocation is reported instead.

Roadmap

  • Native HamLib/OpenFermion adapter
  • Explicit mat2qubit and fermionic conversion path
  • CPU MPO construction and QR/SVD compression
  • CPU MPS observables and variance
  • Two-site DMRG solver
  • Dense CPU validation
  • Optional CuPy CUDA MPO application and observables
  • Reproducible CSV/JSON benchmarks and plots
  • Validate and benchmark CUDA on an RTX 4060 Laptop GPU
  • GPU-resident DMRG effective eigensolver
  • Profile-driven grouped/batched GPU contractions
  • Profile-driven Triton kernel
  • Optional OpenMP backend
  • Multi-GPU execution
  • Molecular orbital-ordering heuristics

Citation and attribution

HamLib is described by Sawaya et al. in arXiv:2306.13126. Preserve the citations and licenses of HamLib, OpenFermion, and mat2qubit when using their data or methods.

hamlib_functions is maintained by Azulene Labs and NERSC contributors Nicolas Sawaya, Katie Klymko, and Daan Camps. This tensor-network contribution extends their public APIs; it does not imply additional affiliation or endorsement.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages