Equation guides and research

LaTeX to DOCX: Native Equation Structure Audit

ImageToLaTeX research · Version 2026-10-02-v1 · Updated

An equation inside a DOCX can be stored as a native equation object instead of a screenshot. This audit checks what the pinned ImageToLaTeX LaTeX → MathML → OMML exporter emits for a disclosed, authored corpus. All 32 prescribed valid examples met their selected structural checks. The sample DOCX contains 32 native equation objects. These results do not certify visual fidelity, mathematical equivalence or behavior in a Word application.

What “native equation” establishes

OMML is Office Math Markup Language. Microsoft documents its equation formats and LaTeX input in Word, and Pandoc documents OMML for mathematics in Word output. Those references describe the representation; they do not prove that this file opens, renders or stays editable in a particular Word version.

Here, “native” means the DOCX stores m:oMath elements containing selected fraction, radical, script, operator or matrix structures. We inspect that markup and the ZIP package. We did not complete a Word-client, full OOXML-schema, accessibility or visual comparison test.

Results by input class

Authored input classCasesParser acceptedSelected structural checks
Prescribed valid examples3232All 32 met their predeclared tags/text
Invalid-input probes60Reported separately below
Extended syntax probes44All 4 met selected checks; feature fidelity was not assessed

The 32 valid examples cover fractions, roots, scripts, n-ary operators, matrices, alignment, fences, accents, functions, Greek symbols and selected styles. A fraction check asks for an m:f element; a root check asks for m:rad. These checks can detect a missing structure without establishing that every detail of a formula was preserved.

Citable finding. In the disclosed ImageToLaTeX authored corpus, all 32 prescribed valid examples produced the selected expected OMML elements/text. The generated DOCX contains 32 native equation objects in well-formed WordprocessingML; this is a structural result, not a rendered or semantic accuracy score.

Worked structural examples

Case IDInput LaTeXSelected required structureResult
fraction\frac{a}{b}m:fMet
square-root\sqrt{x+1}m:radMet
both-scriptsx_i^2m:sSubSupMet
sum\sum_{i=1}^{n}a_im:nary, m:sSubMet
matrix\begin{matrix}1&2\\3&4\end{matrix}m:m, m:mrMet
aligned\begin{aligned}a&=b\\c&=d\end{aligned}m:eqArrMet
greek\alpha+\beta=\gammatext tokensMet

A Met result refers only to the selected predeclared check. The downloadable XML includes the LaTeX and case label beside each equation for independent inspection.

Parser and exporter boundaries

The six invalid probes include an unclosed fraction, unknown command, empty input, whitespace, unclosed matrix and extra brace. The parser rejects all six. Calling the lower-level OMML function directly still returns an empty equation for empty and whitespace probes; the other four calls fail.

Citable finding. The pinned parser rejects all six invalid probes, but direct calls to the lower-level exporter produce empty OMML for blank input. Parser acceptance and emitted markup must be reported separately; this does not establish that the public UI accepts blank formulas.

The four extended probes exercise cancellation, color, a custom command and spacing. Selected structural/text checks pass, but this rubric does not measure whether color, cancellation presentation or spacing survives. MathJax’s differences from TeX explain why a formula parser is not a full TeX document system.

Across all 42 cases, 36 are parser-accepted and 38 produce markup that meets the chosen checks. 38/42 is not a conversion-accuracy rate: it mixes valid formulas, extended probes and two empty objects.

Method and disclosed limits

We authored the 42 cases and their selected expected tags/text before execution. We ran the site’s pinned MathJax → MathML → OMML pipeline, recorded parser and exporter outcomes independently, and generated a DOCX containing the 32 valid cases. An independent Python check validated ZIP integrity, parsed word/document.xml, counted 32 native objects and checked the frozen document SHA-256.

This is first-party testing of our own exporter. No OCR, paid provider call, competitor, customer upload or real document corpus was used. It cannot establish OCR accuracy, arbitrary .tex document performance, cross-platform compatibility or overall mathematical fidelity.

Downloads and reproduction

Download the frozen corpus, results, sample DOCX and verification materials together. Each file is part of the dated audit package; use the case IDs and the version above when comparing a result or reporting a correction.

Save the corpus, results, DOCX, XML and Python script together; run python3 verify-docx.py to inspect the frozen package with the standard library. The exporter rerun requires an existing checkout with exact module hashes and dependencies in the pins manifest; regenerated DOCX byte hashes can differ because package timestamps differ.

Use the audit with the right next step

For formula syntax, see the LaTeX in Word guide and the Word equation editor guide. For a whole PDF, use PDF to Word; review plans and credit prices before a paid document conversion. This audit does not benchmark those separate paths.

Report a correction with a case ID