Skip to content

Jena 6 2 0 - #446

Open
svanteschubert wants to merge 2 commits into
masterfrom
jena-6-2-0
Open

svanteschubert wants to merge 2 commits into
masterfrom
jena-6-2-0

Conversation

@svanteschubert

Copy link
Copy Markdown
Contributor

@xzel23, what do you think? Would you approve the PR?

@Horcrux7, as you requested, a release to update Jena:

Update to Apache Jena 6.2.0 and JDK 21, and replace the java-rdfa dependency

Refs #445

Why

  • Jena 6 requires Java 21. The Jena 6.2.0 jars are compiled for Java 21 (class file version 65), so updating Jena means moving the build and CI from JDK 17 to JDK 21. The released 0.13.0 is not affected and still runs on Java 11.
  • Jena 6 no longer provides jena-iri. Our RDFa parser net.rootdev:java-rdfa:1.0.0-BETA1 still needs that library. Without a fix, Maven quietly falls back to java-rdfa's own transitive jena-iri 3.16.0 from 2020. Pinning jena-iri 5.6.0 would work for now, but that version will get no further updates.
  • java-rdfa is no longer maintained, and there is no maintained replacement: Jena has no RDFa parser, Semargl has been inactive for years, and Apache Any23 was retired in 2023.

Why ODF doesn't need an RDFa library

ODF uses only a small, self-contained subset of RDFa 1.0 (ODF 1.2 Part 1, §4.2.1 and §19.905–§19.908):

  • four attributes: xhtml:about, xhtml:property, xhtml:datatype and xhtml:content
  • on six elements: text:p, text:h, text:meta, text:bookmark-start, table:table-cell and table:covered-table-cell
  • the schema only allows xhtml:about and xhtml:property together

So there is no rel/rev/typeof, and no subject inheritance between nested elements. Each element states its own statements: one per predicate, with the object taken from xhtml:content or the element's text.

ODFDOM already used a modified copy of the java-rdfa parser (about 1,300 lines in 9 classes) that ran as a second SAX handler on every XML file load. This PR replaces all of that with one class, InContentMetadata (about 60 lines of code, with detailed Javadoc).

What changed

  • Dependencies: Jena jena-core and jena-arq 6.2.0 (ARQ is needed for RDF/XML in Jena 6). Removed java-rdfa, the jena-iri pin, and commons-validator, which only the removed code used.
  • JDK 21 in the POM, the CI workflows (maven.yml, deployment.yml), the README and the Eclipse project settings.
  • New org.odftoolkit.odfdom.pkg.rdfa.InContentMetadata: reads the RDF statements of one element, expanding CURIEs through the XML namespace declarations of the file. Its Javadoc documents the rules, with spec references.
  • New package-info.java: explains the two ways ODFDOM reads in-content metadata (the GRDDL stylesheet on the saved files, or the cache on the DOM), how the cache is kept up to date, and the history of this change.
  • The cache is built lazily from the DOM the first time it is requested. Loading a document no longer runs an RDFa parser.
  • BookmarkRDFMetadataExtractor uses the same code for bookmark metadata.

API impact

Unchanged: all methods that return Jena models, i.e. getRDFMetadata(), getInContentMetadata(), getManifestRDFMetadata(), getInContentMetadataFromCache(), getBookmarkRDFMetadata(), OdfFileDom.getInContentMetadataCache(), OdfFileDom.updateInContentMetadataCache(Node) and BookmarkRDFMetadataExtractor. The GRDDL-based getInContentMetadata() / getRDFMetadata() never used java-rdfa.

Removed: these were public only because pkg.rdfa is an exported package, and nothing else in the repository used them:

  • JenaSink, SAXRDFaParser, DOMRDFaParser, URIExtractor, DOMAttributes, MultiContentHandler
  • OdfFileDom.getSink(), whose Javadoc said "The end users needn't to care of this method"
  • OdfFileSaxHandler.setSink(JenaSink)

Moving off them is simple: read metadata with OdfFileDom.getInContentMetadataCache(), or with InContentMetadata.addStatements(model, element, text) for a single element. OdfFileDom.getRDFBaseUri() is new and takes the place of what getSink() was used for.

Note that Jena's Model type is part of the ODFDOM API, so moving to Jena 6 is already a breaking change for users who depend on Jena directly. That makes this release a good point to remove the internal classes too.

Behaviour changes

  • Bookmark literals no longer keep leading and trailing whitespace. Cache literals already worked this way.
  • Bookmarks now support several properties in xhtml:property and support xhtml:datatype. Before, a list of properties was expanded as if it were a single CURIE.
  • Blank-node subjects ([_:x]) now also work when the cache is updated after loading. Before, that path would throw a NullPointerException.
  • Literals are always plain or typed literals, never XML literals, because ODF defines the object as the element's "literal content". xml:lang is not read, since the ODF schema does not allow it on these elements.

Unchanged on purpose: relative IRIs in xhtml:about are still used as they are (the GRDDL stylesheet resolves #id against the file's IRI instead).

Testing

  • Full odfdom build: 629 tests, 0 failures (21 skipped, already marked @Ignore), plus 2 integration tests.
  • New RDFMetadataTest.testGetInContentMetadataFromCache covers the cache: safe CURIE subjects, several predicates, typed literals, xhtml:content taking precedence over the text, updating the text, and removing elements. The previous test for the cache had been commented out.
  • New RDFMetadataTest.testRdfXmlRoundTrip checks that RDF/XML reading and writing still works with Jena 6 (needs ARQ).
  • Old and new output compared on test_rdfmeta.odt: the cache triples are identical. The only difference is the bookmark whitespace trimming described above.

Is it worth it?

Yes:

  • We get off an unmaintained RDFa library and a library Jena no longer ships.
  • About 1,600 lines of code are removed, and the replacement is small, readable and documented.
  • Documents load faster, because metadata is only processed when someone asks for it.
  • The only API break is internal classes that nothing used, and moving off them is straightforward.

…#445)

ODF uses only a self-contained subset of RDFa, so InContentMetadata now builds the per-element RDF cache lazily from the DOM instead of running the unmaintained java-rdfa parser on every XML load. This drops java-rdfa, its legacy jena-iri pin and commons-validator. The internal RDFa plumbing classes of org.odftoolkit.odfdom.pkg.rdfa are removed; the Jena model API is unchanged.
Align odfdom JDT compiler settings with Java 21 and refresh the resource filters written by the Java language server.
@sonarqubecloud

Copy link
Copy Markdown

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants