Archives, libraries, and document collections are rarely flat files: documents have internal structure -- pages, sections, regions, entities -- and rich relationships to other documents, people, places, and events. A Document Representation Model (DRM) captures this structure as a graph, where nodes stand for the objects that make up (or are described by) a document, and typed edges capture how those objects relate to, contain, or depend on one another. This graph-first view makes document content queryable, composable, and reusable across archival, historical, and document-analysis workflows.
CVCDocDB is a Python library developed by the Document Analysis Group (DAG) at the Centre de Visió per Computador (CVC), within the framework of the SUKIDI project, to represent the contents of documents according to a Document Representation Model. It offers a graph-based API with two interchangeable backends -- a persistent Neo4j store and an in-memory NetworkX store for testing and tutorials -- together with semantic entity definitions, WeakNode hierarchies with cascade delete, foreign key validation, vector search, and ready-to-run example datasets for getting started quickly.
- Two backends: Full Neo4j integration (
Neo4jGraph, targeting Neo4j Community Edition; opt-in Enterprise-only features, see Neo4j editions) a Memgraph backend (MemgraphGraph, same propagation policy as Neo4j), an Apache Jena SPARQL backend (JenaGraph, RDF 1.2 in Fuseki), or in-memory NetworkX (NetworkXGraph) for testing and tutorials - WeakNode hierarchy: Child entities with composite primary keys and automatic cascade delete through parent-child edges. Inserting a WeakNode inserts all its ancestors. A chain has at most
MAX_WEAK_CHAIN_DEPTH= 3 nodes (the root plus two levels, e.g.Document → Section → Page). Deeper WeakNodes emit aWeakNodeDepthWarning, and the next major version will give them an automatic surrogate key - ON DELETE strategies: CASCADE, RESTRICT, SET NULL -- choose the deletion semantics that fit your use case
- Semantic entities (optional module, see below): Domain-specific node types such as
IndividuPadro,LlocPadro, andFotografia - FK validation: Foreign key constraints on relations prevent dangling references
- Query and filtering: Secondary index on scalar properties, multi-filter search with intersection/union, debug snapshots
- Vector search (NetworkX only): HNSW-based ANN indexing on node properties with
cosine,l2, andipdistance spaces - RDF/OWL ontology conversion: Generate Python entity classes from RDF/OWL ontologies (RiC-O, etc.)
Install from PyPI:
pip install cvcdocdbOptional features are installed as extras, e.g. pip install "cvcdocdb[rdf,vector]":
| Extra | Enables |
|---|---|
rdf |
RDF/OWL ontology import (cvcdocdb.rdf_schema) |
schema |
Entity-class generation from a YAML schema (cvcdocdb.schema_gen) |
vector |
Vector indexes on NetworkXGraph |
torch |
PyTorch / PyG data loaders (cvcdocdb.torch_dataloader) |
graphrag |
Text2Cypher on a Neo4jGraph (Python >= 3.10). Not needed on MemgraphGraph or NetworkXGraph |
Or install from source in development mode (requirements.txt adds the
documentation/notebook tools and some optional dependencies used for development):
git clone https://github.com/CVC-DAG/cvcdocdb.git
cd cvcdocdb
pip install -e . -r requirements.txtRegister the recommended Jupyter kernel for tutorials:
python -m ipykernel install --user --name cvcdocdb --display-name "Python (cvcdocdb)"from cvcdocdb import NetworkXGraph, Node, WeakNode
# In-memory backend -- no database required
graph = NetworkXGraph()
# Create a document hierarchy
doc = Node(pk={"doc": "DOC-001"}, main_label="Document")
graph.insertNode(doc)
section = WeakNode(parent=doc, pk={"section": 1}, main_label="Section")
graph.insertNode(section, insert_parent=True)
page = WeakNode(parent=section, pk={"page": 1}, main_label="Page")
graph.insertNode(page, insert_parent=True)
# Query the graph
print("Nodes:", graph.get_node_ids())
print("Edges:", graph.get_edges())
graph.close()The core (Node/WeakNode/Relation, GraphStore, migrate() and the
backends) is what import cvcdocdb loads. These modules build on it and are
optional: they aren't imported by import cvcdocdb, and some need an extra.
| Module | Install | What for |
|---|---|---|
cvcdocdb.drm_entities |
(none) | DRM semantic entities: IndividuPadro, LlocPadro, Fotografia... |
cvcdocdb.rico_entities |
(none) | RiC-O archival entities generated from the ontology |
cvcdocdb.rdf_schema |
[rdf] |
RDF/OWL ontology → YAML schema → entity classes |
cvcdocdb.schema_gen |
[schema] |
Entity classes from a YAML schema |
cvcdocdb.torch_dataloader |
[torch] |
Stream a graph into PyTorch / PyTorch Geometric |
cvcdocdb.text2query |
as the translator | Natural-language questions on any backend (SPARQL on Jena, Cypher elsewhere) |
cvcdocdb.text2cypher |
[graphrag], on Neo4j only |
Natural-language questions → Cypher |
cvcdocdb.text2sparql |
(none) | Natural-language questions → SPARQL (JenaGraph) |
from cvcdocdb.drm_entities import IndividuPadro # not: from cvcdocdb import IndividuPadrofrom cvcdocdb import IndividuPadro (and the other DRM entities) still works
but emits a DeprecationWarning; it will be removed in the next major version.
Runnable Jupyter notebooks in docs/tutorials/notebooks/. Each notebook installs the package automatically from the latest release when run.
You can also view them rendered in the hosted documentation.
These use only the common GraphStore API, so they run unchanged on
NetworkXGraph, Neo4jGraph and MemgraphGraph:
intro_basics-- Minimal end-to-end workflow: insert nodes, create WeakNode hierarchiesquerying_and_filtering-- Query operations:get_node(),find_nodes(), property filteringweaknodes_interactive-- Build hierarchies with an interactive widget paneldelete_strategies-- Compare CASCADE, RESTRICT, SET NULL strategies- Datasets, each loaded into NetworkX and Neo4j:
karate_club,movies,game_of_thrones,bibliography_openalex
- NetworkX:
vector_search-- HNSW vector indexing and nearest-neighbor search (NetworkXGraphonly) - Neo4j:
propagation_demo-- full propagation workflow on a real Neo4j database (also as a script:python -m cvcdocdb.exemples.demo_propagation) - Memgraph: no specific notebook; every general one runs on
MemgraphGraphas is
cvcdocdb.rdf_schema/cvcdocdb.schema_gen:generating_classes_from_owl-- Generate Python entity classes from RDF/OWL ontologiescvcdocdb.rico_entities:ric_o_demo(Neo4j) andric_o_networkx_demo(NetworkX) -- the RiC-O model on each backendcvcdocdb.torch_dataloader:torch_dataloader_bibliography-- PyTorch/PyG dataloader, MetaPath2Vec training and link prediction
Generate Python entity classes from RDF/OWL ontologies in one step:
from cvcdocdb.rdf_schema import download_ontology_and_convert
# Downloads, converts to YAML, and generates Python classes
output_path = download_ontology_and_convert(
"https://raw.githubusercontent.com/ICA-EGAD/RiC-O/master/ontology/current-version/RiC-O_1-1.rdf",
"rico",
output_dir="cvcdocdb/"
)
# Generates cvcdocdb/rico_entities.py (677 classes from RiC-O)Step by step:
from cvcdocdb.rdf_schema import download_ontology, rdf_to_yaml
from cvcdocdb.schema_gen import generate_classes
# 1. Download ontology
ont_path = download_ontology(url, output_dir="ontologies/")
# 2. Convert RDF to YAML DRM
yaml_str = rdf_to_yaml(ont_path, "my_db")
# 3. Generate Python classes
py_source = generate_classes(yaml_str)
# 4. Write file
with open("cvcdocdb/entities_my_db.py", "w") as f:
f.write(py_source)The pipeline maps OWL constructs to DRM:
owl:Class-- Node labelrdfs:subClassOf--WeakNodehierarchy (parent)owl:DatatypeProperty-- Node propertiesowl:ObjectProperty-- Relationshipsowl:hasKey-- Primary key fieldsrdfs:comment-- Class docstring
The package includes ready-to-run loaders for common graph domains. They
accept any backend (NetworkXGraph, Neo4jGraph, MemgraphGraph); the
neo4j_*/networkx_* module names are historical:
cvcdocdb.exemples.networkx_karate-- Karate Club graph (NetworkX classic)cvcdocdb.exemples.networkx_bibliografia-- Bibliographic references from OpenAlexcvcdocdb.exemples.neo4j_movies-- Movie-domain graphcvcdocdb.exemples.neo4j_got-- Game of Thrones character-house graph
python -m cvcdocdb.exemples --dataset karate --backend networkx
python -m cvcdocdb.exemples --dataset all --backend both --quietfrom cvcdocdb import NetworkXGraph
from cvcdocdb.exemples import load_karate_club, load_bibliografia_openalex
graph = NetworkXGraph()
print(load_karate_club(graph))
print(load_bibliografia_openalex(graph, query="graph database", per_page=15))
graph.close()MemgraphGraph stores the graph in Memgraph. It
subclasses Neo4jGraph (Memgraph speaks Bolt and Cypher, and uses the same
neo4j driver), so the API and the change-propagation policy are
identical to the Neo4j backend:
- Insert: a WeakNode inserts its parent, and the parent→child edge
carries
_propagate=TRUE. A missing parent, child keys that don't match the parent's, or a duplicate key are refused. Dependencies becomeValornodes. - Update:
update=Truemerges attributes.replace=Truedeletes the old node with propagation (its WeakNode descendants too) and recreates it. - Delete: RESTRICT by default,
propagation=Truefor WeakNode descendants,detach=True(CASCADE), oron_delete="set_null". - Relations: FK validation of both endpoints, plus
update/replace.
from cvcdocdb import MemgraphGraph
graph = MemgraphGraph("bolt://localhost:7687", "", "") # Memgraph has no auth by defaulttest/test_memgraph_graph.py runs the same propagation scenarios on Neo4j and
Memgraph and checks that both leave the same graph and raise the same errors.
The only Memgraph-specific code is pk indexes (SHOW INDEX INFO,
CREATE INDEX ON :Label(props)) and label/relationship-type listing for
schema_yaml(). The Neo4j Enterprise features are not available on Memgraph
(EnterpriseFeatureError), and neither is drop_constraint(). Tested with
Memgraph 3.13 (Community).
JenaGraph stores the graph as RDF 1.2 in Apache Jena Fuseki
(version 6.2.0 or later), with the same API and behaviour as every other
backend:
from cvcdocdb import JenaGraph, Node
graph = JenaGraph("http://localhost:3030/ds") # a Fuseki dataset
graph.insertNode(Node(pk={"doc": "D1"}, main_label="Document", title="Padró"))
graph.query("PREFIX label: <urn:cvcdocdb:label/> PREFIX prop: <urn:cvcdocdb:prop/> "
"SELECT ?t WHERE { ?d a label:Document ; prop:title ?t }") # [{'t': 'Padró'}]- Same behaviour: it reuses the
NetworkXGraphlogic, which is identical toNeo4jGraph's, on an in-memory copy of the graph. The graph has to fit in RAM. - Atomic writes: every mutating call, or a whole
batch(), is sent as one atomic SPARQL Update with only what changed. - Concurrent writes: a version stamp detects writes from other clients
and raises
ConcurrentModificationError; nothing is applied. - Querying:
query()runs SPARQL strings on the server. SPARQL Update is refused, so writes always go through the API. Dict filters and Cypher run on the copy. - RDF layout: nodes are
<urn:cvcdocdb:node/ID>, witha label:…andprop:…values. Relationships arerel:…triples, and their properties are RDF 1.2 annotations (?a rel:X ?b {| prop:p ?v |}). - Isolation:
namespaceandgraph_irikeep it apart from other data in the dataset. - Natural language:
Text2SPARQL, orText2Queryfor any backend.
Requires Apache Jena Fuseki >= 6.2.0, checked on connection through /$/server (FusekiVersionError otherwise; pass server_url= if Fuseki is behind a proxy, or check_fuseki_version=False to use another SPARQL 1.2 store at your own risk). Vector indexes are not supported (as on Neo4j).
CVCDocDB targets Neo4j Community Edition. It is the edition the test suite
and CI run on (neo4j:5-community), and Neo4jGraph uses it by default
(edition="community"). Everything described above works on both Community
and Enterprise.
A few Neo4j Enterprise-only features are available as an opt-in mode:
graph = Neo4jGraph(url, user, password, edition="enterprise") # emits a UserWarning
graph.create_node_key_constraint("Document", ["doc"]) # NODE KEY
graph.create_property_existence_constraint("Document", "title") # IS NOT NULL (nodes or relationships)
graph.create_property_type_constraint("Document", "year", "INTEGER") # IS :: TYPE (Neo4j 5.9+)
graph.create_database("projecte1") # multiple databases
graph.drop_database("projecte1")In the default Community mode these methods raise EnterpriseFeatureError
without contacting the server. In Enterprise mode, every call also checks that
the server really is Enterprise (graph.server_edition()).
Warning: the Enterprise features require a valid Neo4j Enterprise license and are not fully tested. The regular suite only checks the generated Cypher and the Community-mode guards. The tests against a real Enterprise server only run when
NEO4J_ENTERPRISE_URLis set, and CI does not set it. No constraint is created automatically. ANODE KEYon a label used with several pk shapes rejects the nodes of the other shapes.
CVCDocDB uses environment variables for Neo4j connections. Multiple targets are supported via the NEO4J_TARGET selector:
# Default target
export NEO4J_DEV_URL=bolt://dev-host:7687
export NEO4J_DEV_USER=neo4j
export NEO4J_DEV_PASSWORD=your_dev_password
export NEO4J_DEV_DATABASE=neo4j
# Custom target
export NEO4J_TARGET=LOCAL
export NEO4J_LOCAL_URL=bolt://localhost:7687
export NEO4J_LOCAL_USER=neo4j
export NEO4J_LOCAL_PASSWORD=your_password
export NEO4J_LOCAL_DATABASE=neo4jpip install -r requirements-test.txt
python -m pytest test/ -vThree test levels:
- Unit (
-m unit) -- fast, no graph store - Integration (
-m integration) -- NetworkXGraph (in-memory) - Neo4j (
-m slow) -- requires a real Neo4j connection. If none is reachable and Docker is available, a disposableneo4j:5-communitycontainer is started automatically; otherwise these tests auto-skip. Seetest/README.mdfor details and the manualdocker-compose.neo4j.ymloption. - Memgraph (
-m slow,test/test_memgraph_graph.py) -- requires a Memgraph server atMEMGRAPH_URL(MEMGRAPH_USER/MEMGRAPH_PASSWORD). If none is reachable and Docker is available, a disposablememgraph/memgraph:3.13.1container is started on port 7688. - Apache Jena (
-m slow,test/test_jena_graph.py) -- requires a Fuseki dataset atFUSEKI_URL(e.g.http://localhost:3030/ds). If none is reachable and Docker is available, a disposablesecoresearch/fuseki:6.2.0container is started on port 3030.
Skip Neo4j tests: pytest test/ -v -m "not slow"
- Hosted docs: https://cvc-dag.github.io/cvcdocdb/
- Source docs:
docs/-- Sphinx documentation source
Generate HTML docs with Sphinx:
cd docs
sphinx-build -b html . _build/htmlCVCDocDB itself is licensed under the GNU General Public License v3 or later (see LICENSE). It does not include or redistribute any database server. It talks to Neo4j, Memgraph and Apache Jena Fuseki over their network protocols (Bolt, HTTP/SPARQL), and each server is a separate product under its own license. Installing, running and licensing a server is the responsibility of whoever deploys it.
| Software | Used by | License |
|---|---|---|
| Neo4j Community Edition | Neo4jGraph (default) |
GPL-3.0 |
| Neo4j Enterprise Edition | Neo4jGraph(edition="enterprise") |
Commercial: needs an Enterprise license |
| Memgraph | MemgraphGraph |
Business Source License 1.1. It is source-available, not open source: production use is limited by its Additional Use Grant, and it converts to Apache-2.0 on its Change Date. Enterprise features fall under the Memgraph Enterprise License. |
| Apache Jena Fuseki | JenaGraph |
Apache-2.0 |
About Memgraph's license. Memgraph Community Edition is under the Memgraph Business Source License 1.1. It lets you copy, modify and use Memgraph for non-production purposes, and, through its Additional Use Grant, in production for your own internal purposes. It does not allow you to:
- embed or distribute Memgraph to third parties, or give third parties direct access to operate or control it as a standalone solution or service;
- offer it as a database-as-a-service (or any equivalent model);
- build a product that competes with Memgraph.
Using Memgraph with CVCDocDB inside your organisation (e.g. a research group's own projects and users) is therefore internal use. Offering Memgraph itself to third parties needs a commercial license from Memgraph. Each Memgraph version becomes Apache-2.0 on the license's Change Date, or four years after that version's first release, whichever comes first. Enterprise features (e.g. multi-tenancy) need a Memgraph Enterprise License.
Python packages installed with CVCDocDB are under licenses compatible with the GPL-3.0:
neo4j(driver): Apache-2.0networkx,numpy: BSD-3-Clausefilelock: MITtqdm: MPL-2.0 and MIT- Optional extras:
neo4j-graphragandhnswlib(Apache-2.0),rdflib(BSD-3-Clause),pyyamlandtorch_geometric(MIT),torch(BSD-style; see its metadata).
The Docker images used by the test suite and CI (neo4j:5-community,
memgraph/memgraph, secoresearch/fuseki) are only pulled to run tests. They
are not part of the CVCDocDB package.
This section is informative, not legal advice: check each license for your use case.
- Oriol Ramos Terrades
- Jialuo Chen
- Adrià Molina
This work has been partially supported by the Spanish projects PID2021-126808OB-I00 and PID2024-157778OB-I00, Ministerio de Ciencia e Innovación, the Departament de Cultura of the Generalitat de Catalunya, and the CERCA Program / Generalitat de Catalunya. Adrià Molina is funded with the PRE2022-101575 grant provided by MCIN / AEI / 10.13039 / 501100011033 and by the European Social Fund (FSE+).