Update, September 2026The answer above was right in 2012, but it isn't any more, so for anyone landing here from a search:
docx4j now has a layout model, courtesy of Apache FOP. Two things changed:
- 17.1.0: the PDF/XSL-FO output (docx4j-export-fo) was reworked to follow Word's layout rules by default - line breaking, paragraph spacing, keeps, table rows, page breaks and so on - measured against Word's own page counts over a corpus of real documents. So FOP's layout of a docx4j FO rendering is now a reasonable stand-in for Word's.
- 17.1.1: a new class org.docx4j.model.pagination.Paginate reads FOP's area tree (its record of what it laid out where) and turns it into data about your document.
What Paginate gives you directly:
- Code: Select all
PaginationMap map = Paginate.compute(wordMLPackage, null);
map.getPage(paraId); // the page a paragraph starts on
map.getBreaks(paraId); // for a paragraph that spans pages, the character offsets where each new page begins
map.getPageCount();
Paragraphs are keyed by their w14:paraId (assigned if missing).
Paginate.paginate(...) additionally writes the result back into the document as w:lastRenderedPageBreak markers, exactly as Word records its own last rendering. The TOC generator's page numbers now come from the same place.
For the original question -
how many lines a paragraph occupies, and where each line breaks - Paginate doesn't expose that, but the information is there. FOP's area tree contains one
lineArea element for every line it laid out, with the text of that line, inside a block per paragraph. To get at it, ask for the area tree instead of a PDF:
- Code: Select all
FOSettings foSettings = Docx4J.createFOSettings();
foSettings.setOpcPackage(wordMLPackage);
foSettings.setApacheFopMime("application/X-fop-areatree");
foSettings.getFeatures().add(ConversionFeatures.PP_FO_PARAGRAPH_IDS); // each block gets prod-id="p-" + paraId
Docx4J.toFO(foSettings, outputStream, Docx4J.FLAG_EXPORT_PREFER_NONXSL);
Then walk the resulting XML: each
block with a prod-id is one of your paragraphs, and its
lineArea children are its lines. Paginate's own reader (
PaginationAreaTreeHandler) is a worked example of parsing it.
Two caveats. It is FOP's layout, using the fonts FOP can see, so it is close to Word rather than identical (the closer your fonts match, the closer the result). And it needs docx4j-export-fo on the classpath, since the layout is a real rendering.