Skip to content
CVC-DAGPublic

About

dlm is a package for document language model representation

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

Document Representation Models

Archives, libraries, and document collections are rarely flat files: documents have internal structure -- pages, sections, regions, entities -- and rich relationships to other documents, people, places, and events. A Document Representation Model (DRM) captures this structure as a graph, where nodes stand for the objects that make up (or are described by) a document, and typed edges capture how those objects relate to, contain, or depend on one another. This graph-first view makes document content queryable, composable, and reusable across archival, historical, and document-analysis workflows.

CVCDocDB

CVCDocDB is a Python library developed by the Document Analysis Group (DAG) at the Centre de Visió per Computador (CVC), within the framework of the SUKIDI project, to represent the contents of documents according to a Document Representation Model. It offers a graph-based API with two interchangeable backends -- a persistent Neo4j store and an in-memory NetworkX store for testing and tutorials -- together with semantic entity definitions, WeakNode hierarchies with cascade delete, foreign key validation, vector search, and ready-to-run example datasets for getting started quickly.

Features

  • Two backends: Full Neo4j integration (Neo4jGraph, targeting Neo4j Community Edition; opt-in Enterprise-only features, see Neo4j editions) a Memgraph backend (MemgraphGraph, same propagation policy as Neo4j), an Apache Jena SPARQL backend (JenaGraph, RDF 1.2 in Fuseki), or in-memory NetworkX (NetworkXGraph) for testing and tutorials
  • WeakNode hierarchy: Child entities with composite primary keys and automatic cascade delete through parent-child edges. Inserting a WeakNode inserts all its ancestors. A chain has at most MAX_WEAK_CHAIN_DEPTH = 3 nodes (the root plus two levels, e.g. Document → Section → Page). Deeper WeakNodes emit a WeakNodeDepthWarning, and the next major version will give them an automatic surrogate key
  • ON DELETE strategies: CASCADE, RESTRICT, SET NULL -- choose the deletion semantics that fit your use case
  • Semantic entities (optional module, see below): Domain-specific node types such as IndividuPadro, LlocPadro, and Fotografia
  • FK validation: Foreign key constraints on relations prevent dangling references
  • Query and filtering: Secondary index on scalar properties, multi-filter search with intersection/union, debug snapshots
  • Vector search (NetworkX only): HNSW-based ANN indexing on node properties with cosine, l2, and ip distance spaces
  • RDF/OWL ontology conversion: Generate Python entity classes from RDF/OWL ontologies (RiC-O, etc.)

Installation

Install from PyPI:

pip install cvcdocdb

Optional features are installed as extras, e.g. pip install "cvcdocdb[rdf,vector]":

Extra Enables
rdf RDF/OWL ontology import (cvcdocdb.rdf_schema)
schema Entity-class generation from a YAML schema (cvcdocdb.schema_gen)
vector Vector indexes on NetworkXGraph
torch PyTorch / PyG data loaders (cvcdocdb.torch_dataloader)
graphrag Text2Cypher on a Neo4jGraph (Python >= 3.10). Not needed on MemgraphGraph or NetworkXGraph

Or install from source in development mode (requirements.txt adds the documentation/notebook tools and some optional dependencies used for development):

git clone https://github.com/CVC-DAG/cvcdocdb.git
cd cvcdocdb
pip install -e . -r requirements.txt

Register the recommended Jupyter kernel for tutorials:

python -m ipykernel install --user --name cvcdocdb --display-name "Python (cvcdocdb)"

Quick Start

from cvcdocdb import NetworkXGraph, Node, WeakNode

# In-memory backend -- no database required
graph = NetworkXGraph()

# Create a document hierarchy
doc = Node(pk={"doc": "DOC-001"}, main_label="Document")
graph.insertNode(doc)

section = WeakNode(parent=doc, pk={"section": 1}, main_label="Section")
graph.insertNode(section, insert_parent=True)

page = WeakNode(parent=section, pk={"page": 1}, main_label="Page")
graph.insertNode(page, insert_parent=True)

# Query the graph
print("Nodes:", graph.get_node_ids())
print("Edges:", graph.get_edges())
graph.close()

Optional modules

The core (Node/WeakNode/Relation, GraphStore, migrate() and the backends) is what import cvcdocdb loads. These modules build on it and are optional: they aren't imported by import cvcdocdb, and some need an extra.

Module Install What for
cvcdocdb.drm_entities (none) DRM semantic entities: IndividuPadro, LlocPadro, Fotografia...
cvcdocdb.rico_entities (none) RiC-O archival entities generated from the ontology
cvcdocdb.rdf_schema [rdf] RDF/OWL ontology → YAML schema → entity classes
cvcdocdb.schema_gen [schema] Entity classes from a YAML schema
cvcdocdb.torch_dataloader [torch] Stream a graph into PyTorch / PyTorch Geometric
cvcdocdb.text2query as the translator Natural-language questions on any backend (SPARQL on Jena, Cypher elsewhere)
cvcdocdb.text2cypher [graphrag], on Neo4j only Natural-language questions → Cypher
cvcdocdb.text2sparql (none) Natural-language questions → SPARQL (JenaGraph)
from cvcdocdb.drm_entities import IndividuPadro   # not: from cvcdocdb import IndividuPadro

from cvcdocdb import IndividuPadro (and the other DRM entities) still works but emits a DeprecationWarning; it will be removed in the next major version.

Tutorial Notebooks

Runnable Jupyter notebooks in docs/tutorials/notebooks/. Each notebook installs the package automatically from the latest release when run.

You can also view them rendered in the hosted documentation.

General (any backend)

These use only the common GraphStore API, so they run unchanged on NetworkXGraph, Neo4jGraph and MemgraphGraph:

  • intro_basics -- Minimal end-to-end workflow: insert nodes, create WeakNode hierarchies
  • querying_and_filtering -- Query operations: get_node(), find_nodes(), property filtering
  • weaknodes_interactive -- Build hierarchies with an interactive widget panel
  • delete_strategies -- Compare CASCADE, RESTRICT, SET NULL strategies
  • Datasets, each loaded into NetworkX and Neo4j: karate_club, movies, game_of_thrones, bibliography_openalex

Backend-specific

  • NetworkX: vector_search -- HNSW vector indexing and nearest-neighbor search (NetworkXGraph only)
  • Neo4j: propagation_demo -- full propagation workflow on a real Neo4j database (also as a script: python -m cvcdocdb.exemples.demo_propagation)
  • Memgraph: no specific notebook; every general one runs on MemgraphGraph as is

Optional modules

  • cvcdocdb.rdf_schema / cvcdocdb.schema_gen: generating_classes_from_owl -- Generate Python entity classes from RDF/OWL ontologies
  • cvcdocdb.rico_entities: ric_o_demo (Neo4j) and ric_o_networkx_demo (NetworkX) -- the RiC-O model on each backend
  • cvcdocdb.torch_dataloader: torch_dataloader_bibliography -- PyTorch/PyG dataloader, MetaPath2Vec training and link prediction

RDF/OWL Ontology Conversion

Generate Python entity classes from RDF/OWL ontologies in one step:

from cvcdocdb.rdf_schema import download_ontology_and_convert

# Downloads, converts to YAML, and generates Python classes
output_path = download_ontology_and_convert(
    "https://raw.githubusercontent.com/ICA-EGAD/RiC-O/master/ontology/current-version/RiC-O_1-1.rdf",
    "rico",
    output_dir="cvcdocdb/"
)
# Generates cvcdocdb/rico_entities.py (677 classes from RiC-O)

Step by step:

from cvcdocdb.rdf_schema import download_ontology, rdf_to_yaml
from cvcdocdb.schema_gen import generate_classes

# 1. Download ontology
ont_path = download_ontology(url, output_dir="ontologies/")

# 2. Convert RDF to YAML DRM
yaml_str = rdf_to_yaml(ont_path, "my_db")

# 3. Generate Python classes
py_source = generate_classes(yaml_str)

# 4. Write file
with open("cvcdocdb/entities_my_db.py", "w") as f:
    f.write(py_source)

The pipeline maps OWL constructs to DRM:

  • owl:Class -- Node label
  • rdfs:subClassOf -- WeakNode hierarchy (parent)
  • owl:DatatypeProperty -- Node properties
  • owl:ObjectProperty -- Relationships
  • owl:hasKey -- Primary key fields
  • rdfs:comment -- Class docstring

Example Dataset Loaders (cvcdocdb.exemples)

The package includes ready-to-run loaders for common graph domains. They accept any backend (NetworkXGraph, Neo4jGraph, MemgraphGraph); the neo4j_*/networkx_* module names are historical:

  • cvcdocdb.exemples.networkx_karate -- Karate Club graph (NetworkX classic)
  • cvcdocdb.exemples.networkx_bibliografia -- Bibliographic references from OpenAlex
  • cvcdocdb.exemples.neo4j_movies -- Movie-domain graph
  • cvcdocdb.exemples.neo4j_got -- Game of Thrones character-house graph

Command-line loader

python -m cvcdocdb.exemples --dataset karate --backend networkx
python -m cvcdocdb.exemples --dataset all --backend both --quiet

Programmatic usage

from cvcdocdb import NetworkXGraph
from cvcdocdb.exemples import load_karate_club, load_bibliografia_openalex

graph = NetworkXGraph()
print(load_karate_club(graph))
print(load_bibliografia_openalex(graph, query="graph database", per_page=15))
graph.close()

Memgraph backend

MemgraphGraph stores the graph in Memgraph. It subclasses Neo4jGraph (Memgraph speaks Bolt and Cypher, and uses the same neo4j driver), so the API and the change-propagation policy are identical to the Neo4j backend:

  • Insert: a WeakNode inserts its parent, and the parent→child edge carries _propagate=TRUE. A missing parent, child keys that don't match the parent's, or a duplicate key are refused. Dependencies become Valor nodes.
  • Update: update=True merges attributes. replace=True deletes the old node with propagation (its WeakNode descendants too) and recreates it.
  • Delete: RESTRICT by default, propagation=True for WeakNode descendants, detach=True (CASCADE), or on_delete="set_null".
  • Relations: FK validation of both endpoints, plus update/replace.
from cvcdocdb import MemgraphGraph

graph = MemgraphGraph("bolt://localhost:7687", "", "")  # Memgraph has no auth by default

test/test_memgraph_graph.py runs the same propagation scenarios on Neo4j and Memgraph and checks that both leave the same graph and raise the same errors. The only Memgraph-specific code is pk indexes (SHOW INDEX INFO, CREATE INDEX ON :Label(props)) and label/relationship-type listing for schema_yaml(). The Neo4j Enterprise features are not available on Memgraph (EnterpriseFeatureError), and neither is drop_constraint(). Tested with Memgraph 3.13 (Community).

Apache Jena backend (SPARQL)

JenaGraph stores the graph as RDF 1.2 in Apache Jena Fuseki (version 6.2.0 or later), with the same API and behaviour as every other backend:

from cvcdocdb import JenaGraph, Node

graph = JenaGraph("http://localhost:3030/ds")          # a Fuseki dataset
graph.insertNode(Node(pk={"doc": "D1"}, main_label="Document", title="Padró"))
graph.query("PREFIX label: <urn:cvcdocdb:label/> PREFIX prop: <urn:cvcdocdb:prop/> "
            "SELECT ?t WHERE { ?d a label:Document ; prop:title ?t }")   # [{'t': 'Padró'}]
  • Same behaviour: it reuses the NetworkXGraph logic, which is identical to Neo4jGraph's, on an in-memory copy of the graph. The graph has to fit in RAM.
  • Atomic writes: every mutating call, or a whole batch(), is sent as one atomic SPARQL Update with only what changed.
  • Concurrent writes: a version stamp detects writes from other clients and raises ConcurrentModificationError; nothing is applied.
  • Querying: query() runs SPARQL strings on the server. SPARQL Update is refused, so writes always go through the API. Dict filters and Cypher run on the copy.
  • RDF layout: nodes are <urn:cvcdocdb:node/ID>, with a label:… and prop:… values. Relationships are rel:… triples, and their properties are RDF 1.2 annotations (?a rel:X ?b {| prop:p ?v |}).
  • Isolation: namespace and graph_iri keep it apart from other data in the dataset.
  • Natural language: Text2SPARQL, or Text2Query for any backend.

Requires Apache Jena Fuseki >= 6.2.0, checked on connection through /$/server (FusekiVersionError otherwise; pass server_url= if Fuseki is behind a proxy, or check_fuseki_version=False to use another SPARQL 1.2 store at your own risk). Vector indexes are not supported (as on Neo4j).

Neo4j editions: Community (default) and Enterprise

CVCDocDB targets Neo4j Community Edition. It is the edition the test suite and CI run on (neo4j:5-community), and Neo4jGraph uses it by default (edition="community"). Everything described above works on both Community and Enterprise.

A few Neo4j Enterprise-only features are available as an opt-in mode:

graph = Neo4jGraph(url, user, password, edition="enterprise")  # emits a UserWarning
graph.create_node_key_constraint("Document", ["doc"])           # NODE KEY
graph.create_property_existence_constraint("Document", "title") # IS NOT NULL (nodes or relationships)
graph.create_property_type_constraint("Document", "year", "INTEGER")  # IS :: TYPE (Neo4j 5.9+)
graph.create_database("projecte1")                              # multiple databases
graph.drop_database("projecte1")

In the default Community mode these methods raise EnterpriseFeatureError without contacting the server. In Enterprise mode, every call also checks that the server really is Enterprise (graph.server_edition()).

Warning: the Enterprise features require a valid Neo4j Enterprise license and are not fully tested. The regular suite only checks the generated Cypher and the Community-mode guards. The tests against a real Enterprise server only run when NEO4J_ENTERPRISE_URL is set, and CI does not set it. No constraint is created automatically. A NODE KEY on a label used with several pk shapes rejects the nodes of the other shapes.

Configuration

CVCDocDB uses environment variables for Neo4j connections. Multiple targets are supported via the NEO4J_TARGET selector:

# Default target
export NEO4J_DEV_URL=bolt://dev-host:7687
export NEO4J_DEV_USER=neo4j
export NEO4J_DEV_PASSWORD=your_dev_password
export NEO4J_DEV_DATABASE=neo4j

# Custom target
export NEO4J_TARGET=LOCAL
export NEO4J_LOCAL_URL=bolt://localhost:7687
export NEO4J_LOCAL_USER=neo4j
export NEO4J_LOCAL_PASSWORD=your_password
export NEO4J_LOCAL_DATABASE=neo4j

Running Tests

pip install -r requirements-test.txt
python -m pytest test/ -v

Three test levels:

  • Unit (-m unit) -- fast, no graph store
  • Integration (-m integration) -- NetworkXGraph (in-memory)
  • Neo4j (-m slow) -- requires a real Neo4j connection. If none is reachable and Docker is available, a disposable neo4j:5-community container is started automatically; otherwise these tests auto-skip. See test/README.md for details and the manual docker-compose.neo4j.yml option.
  • Memgraph (-m slow, test/test_memgraph_graph.py) -- requires a Memgraph server at MEMGRAPH_URL (MEMGRAPH_USER/MEMGRAPH_PASSWORD). If none is reachable and Docker is available, a disposable memgraph/memgraph:3.13.1 container is started on port 7688.
  • Apache Jena (-m slow, test/test_jena_graph.py) -- requires a Fuseki dataset at FUSEKI_URL (e.g. http://localhost:3030/ds). If none is reachable and Docker is available, a disposable secoresearch/fuseki:6.2.0 container is started on port 3030.

Skip Neo4j tests: pytest test/ -v -m "not slow"

Documentation

Generate HTML docs with Sphinx:

cd docs
sphinx-build -b html . _build/html

Third-party software and licenses

CVCDocDB itself is licensed under the GNU General Public License v3 or later (see LICENSE). It does not include or redistribute any database server. It talks to Neo4j, Memgraph and Apache Jena Fuseki over their network protocols (Bolt, HTTP/SPARQL), and each server is a separate product under its own license. Installing, running and licensing a server is the responsibility of whoever deploys it.

Software Used by License
Neo4j Community Edition Neo4jGraph (default) GPL-3.0
Neo4j Enterprise Edition Neo4jGraph(edition="enterprise") Commercial: needs an Enterprise license
Memgraph MemgraphGraph Business Source License 1.1. It is source-available, not open source: production use is limited by its Additional Use Grant, and it converts to Apache-2.0 on its Change Date. Enterprise features fall under the Memgraph Enterprise License.
Apache Jena Fuseki JenaGraph Apache-2.0

About Memgraph's license. Memgraph Community Edition is under the Memgraph Business Source License 1.1. It lets you copy, modify and use Memgraph for non-production purposes, and, through its Additional Use Grant, in production for your own internal purposes. It does not allow you to:

  • embed or distribute Memgraph to third parties, or give third parties direct access to operate or control it as a standalone solution or service;
  • offer it as a database-as-a-service (or any equivalent model);
  • build a product that competes with Memgraph.

Using Memgraph with CVCDocDB inside your organisation (e.g. a research group's own projects and users) is therefore internal use. Offering Memgraph itself to third parties needs a commercial license from Memgraph. Each Memgraph version becomes Apache-2.0 on the license's Change Date, or four years after that version's first release, whichever comes first. Enterprise features (e.g. multi-tenancy) need a Memgraph Enterprise License.

Python packages installed with CVCDocDB are under licenses compatible with the GPL-3.0:

  • neo4j (driver): Apache-2.0
  • networkx, numpy: BSD-3-Clause
  • filelock: MIT
  • tqdm: MPL-2.0 and MIT
  • Optional extras: neo4j-graphrag and hnswlib (Apache-2.0), rdflib (BSD-3-Clause), pyyaml and torch_geometric (MIT), torch (BSD-style; see its metadata).

The Docker images used by the test suite and CI (neo4j:5-community, memgraph/memgraph, secoresearch/fuseki) are only pulled to run tests. They are not part of the CVCDocDB package.

This section is informative, not legal advice: check each license for your use case.

Authors and Contributors

  • Oriol Ramos Terrades
  • Jialuo Chen
  • Adrià Molina

Acknowledgements

This work has been partially supported by the Spanish projects PID2021-126808OB-I00 and PID2024-157778OB-I00, Ministerio de Ciencia e Innovación, the Departament de Cultura of the Generalitat de Catalunya, and the CERCA Program / Generalitat de Catalunya. Adrià Molina is funded with the PRE2022-101575 grant provided by MCIN / AEI / 10.13039 / 501100011033 and by the European Social Fund (FSE+).

About

dlm is a package for document language model representation

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages