Clarion Unicode Beta · Companion Document

EMF Page Files — Transition Note for Report Tool Authors

Tech brief for developers who ship report-related products for Clarion: preview replacements, export/document generators, report add-on templates, or anything that reads report page files or implements the report generator interface. What changed, when it appears, and the transition path for each integration style.

Scope third-party report tools Trigger REPORT, …, UNICODE → .emf pages Guarantee non-opted reports: byte-identical WMF

Companion to the Clarion Unicode (USTRING) — Tester Guide. Items marked known limitation are deliberate in this build — not bugs to report.

1. What changed

A report can opt into Unicode with the UNICODE attribute (REPORT,...,UNICODE). Opted-in preview page files are Windows Enhanced Metafiles (.emf) instead of classic placeable WMF. That is the whole delta third-party tools see.

Without the attribute, pages stay byte-identical WMF — same as every prior Clarion release. Since build 14313 the attribute is no longer the only switch, and for this beta release EMF is the default for generated applications (a default we want your feedback on) — plan for install bases that see .emf on every report, not just the ones a developer marked by hand. Three levers turn a report wide, all documented in the tester guide's What's new (Reports and the report previewer):

The same build replaces the shipped previewer: ReportPreviewClass is the default Print Previewer, and an application whose prompt still reads PrintPreviewClass generates the new class on its next generate (a derived or third-party class named there is left alone). Its text index, PageTextIndexClass, is GUI-free and yours to reuse — §8a.

Why EMF: WMF text records are byte-only; they cannot carry UTF-16. EMF uses EXTTEXTOUTW and is the native Windows metafile format. Page model is unchanged: one file per page, same preview queue, same properties — only the on-disk format differs.

2. Format detection

Do not route by filename. Read the header:

Both checks are a few dozen bytes. One header fork covers both formats with no config and no knowledge of whether the report opted in.

Tip
The shipped Image2PDF path used to alias by extension and was the last lane to fail on .emf. It now reads the header. Prefer content detection over the filename.

3. Pick your path

ProductTransitionSection
A. Uses or derives from shipped WMFParser / WMFDocumentParser (ABWMFPAR) Rebuild. §4
B. Implements IReportGenerator (export target) No break. Wide surface shipped, opt-in (IReportGeneratorW) §5
C. Own page-file parser Real work; this note + shipped source are the map §6
D. Displays/prints pages (preview, viewer) Easier than WMF — Windows plays EMF natively §8

Paths compose (often B on A, or D with a slice of C for band navigation). Read every section that applies.

4. Path A — ABWMFPAR consumers

This build's shipped WMFParser already reads both formats. The §2 header check is built in. .emf pages use the same delivery surface (ProcessString, ProcessText, ProcessBand, ProcessImage, fonts, shapes, band/control markers). Coordinates stay in 1/1000-inch space.

Transition: recompile against this build's LibSrc. ABWMFPAR is LibSrc source (not a DLL). Rebuild is the upgrade.

Checks:

After rebuild, text fidelity matches today: wide page text is narrowed to ANSI (best-fit, same as the WMF lane). Wide delivery to generators is the next step — see §5.

5. Path B — IReportGenerator implementors

Generators work unchanged in this build. The shipped parser feeds .emf pages through the same narrow IReportGenerator methods (ProcessString(…, STRING Text, …), ProcessText with CSTRING(257) lines, rest of ABRPTGEN.INT). Wide page text is narrowed before it reaches the interface — best-fit codepage, same as WMF. Recompile and you are done for this build.

Shipped: the wide surface is IReportGeneratorW (Option A2, refined)

The decision landed as A2 — zero break — refined one step further: IReportGeneratorW is a standalone companion interface (it does not inherit IReportGenerator), so adopting means implementing exactly seven methods next to your existing ones:

Everything else — pages, shapes, ordinary images, properties — keeps riding the narrow interface (only the emoji stamps arrive through ProcessImageW). Clarion has no runtime interface query, so registration is an explicit handoff: the narrow ref goes to the parser via Init as always; the wide ref via WMFDocumentParser.SetWideGenerator after Init, or the two-ref selector overload AddItem(gen.IReportGenerator, gen.IReportGeneratorW). The shipped report-output templates do this automatically — a wide-capable target registers both faces exactly when the procedure's REPORT carries UNICODE, and every other combination generates the classic narrow registration byte-identical.

A generator that never adopts keeps compiling and receiving best-fit narrowed text, indefinitely. The shipped TEXT, HTML and XML targets have adopted (UTF-8 output, honest charset/encoding declarations), and so has the PDF target (Type0/CIDFontType2 fonts with glyph subsets and ToUnicode; emoji stamps through ProcessImageW as RGB + soft mask) — see section 7. A narrow-only target fed wide content that actually loses characters posts one “Report Export Notice” per document.

Feedback still useful: would your product adopt the wide surface, and on which targets first?

6. Path C — custom page parsers

The real transition work. Shortcut: this build's LibSrc ABWMFPAR.CLW is a complete working EMF page parser in Clarion — header fork, record walk, comment channel, transforms, images. That file's EMF arms are the reference for everything below.

6a. Closed record dialect

Pages come from the Clarion print engine only — a closed dialect, not "all of EMF":

Mechanical traps vs WMF:

6b. World transforms are live

WMF pages stored final device coordinates in WORD fields (engine already flattened). EMF records under SETWORLDTRANSFORM / MODIFYWORLDTRANSFORM stacked with SAVEDC/RESTOREDC — coordinates in text/shape records are pre-transform. Ignore transform state and geometry is wrong.

Keep a small transform state: current 2×3 matrix, push/pop with the DC stack, apply to every extracted coordinate. The reference does exactly that. Engine usage is simple (scale/translate only; no rotation in report pages), but it must be applied.

6c. Explicit object handles (simpler than WMF)

WMF handles are implicit free-slot allocation. EMF create-records carry an explicit ihObject index — bookkeeping is a straight array. Guard: SELECTOBJECT of a stock object sets ENHMETA_STOCK_OBJECT (high bit) — mask and skip the table lookup or you index garbage.

6d. Band/control comment channel

Band start/end, control start/end (names + text), image names, and the extended-attribute channel on start-control payloads still use the same block structures. Only the carrier changed:

Comment payload starts at byte 13 of the record (1-based; after iType/nSize/cbData) with the same MFCOMMENT magic (0527h). Block layouts are unchanged; fixed offsets shift by the carrier prefix delta. Reference has per-kind offsets.

Tip — GDI merge trap
GDI may merge consecutive GdiComments into one EMR_GDICOMMENT. Walk blocks inside the record, stride by each block's size (fixed kinds fixed; start-control / image-name / comment-info = struct size + embedded string length) until cbData is consumed. Today's engine usually emits one block per record — a one-block parser will pass local tests and fail later when GDI merges. Walk blocks.

6e. Images

EMR_STRETCHDIBITS carries a packed DIB (BITMAPINFO immediately followed by bits). DIB starts at RecordStart + offBmiSrc. Pipelines that re-read image bytes by file offset should publish that offset. bmi+bits contiguity is verified on engine output.

EMR_ALPHABLEND (§6h) carries the same packed-DIB layout: BITMAPINFO at RecordStart + offBmiSrc, bits at RecordStart + offBitsSrc (publish both — the shipped parser passes the bits offset as BITS= to ProcessImageW), 32 bpp BGRA, the BLENDFUNCTION in dwRop (SourceConstantAlpha = byte 2, AlphaFormat = byte 3; AC_SRC_ALPHA = per-pixel, premultiplied). Positive biHeight = bottom-up rows.

6f. Text is UTF-16 (including surrogates)

EXTTEXTOUTW strings are UTF-16LE code units. For store/export/measure/ truncate logic, read the tester guide §2 (UTF-16 code units, BMP vs non-BMP): one unit is not always one character; non-BMP (many emoji, supplementary CJK) is a surrogate pair (two units) — keep pairs intact at any cut. ANSI output: narrow deliberately (best-fit, lossy). UTF-8/UTF-16 output: carry units through.

6g. Coordinate law (the hard part, solved)

Working recipe: derive DPI from the header, snap to nearest multiple of 4 (real drivers are; guard ≤2% deviation), reconstruct physical offset as the centered difference between rclFrame (0.01 mm — ×1000/2540 → 1/1000 inch) and the printable area. Exact for centered printable areas (the norm); bounded approx for asymmetric margins. For page sizing use rclFrame, never rclBounds (tight ink extent in recording pixels — mis-sizes the page).

Tools that only play into a DC (§8) can ignore most of this — transforms resolve themselves. The law matters when extracting geometry (layout, export coords, hit-testing).

6h. Emoji and mixed strings — partitioned text recording

(Beta refresh after 08-10-2026 — pages from earlier beta builds recorded emoji-bearing strings differently; the contract below is the shipped one.)

A string containing no emoji records as one ordinary visible EXTTEXTOUTW — nothing changes. A string that contains emoji records as a partitioned sequence inside its control:

  1. An empty ETO_OPAQUE record carrying the background fill.
  2. The non-emoji segments as visible EXTTEXTOUTW records — real text color, transparent background, UTF-16 units verbatim (any script, not just ASCII).
  3. Each emoji cluster as an EMR_ALPHABLEND color raster stamp only — no text glyph is drawn for it.
  4. One full-string EXTTEXTOUTW record — complete text, real color, original position — bracketed by an empty clip (SAVEDC + INTERSECTCLIPRECT(0,0,0,0) + RESTOREDC), so playback draws nothing from it.

Example: u'Ω中 😀 café' records Ω中  (3 units) + stamp +  café + the clipped 10-unit full string.

What this means per path:

Without DirectWrite on the machine the stamps are absent and emoji record as ordinary (monochrome) text — the pre-refresh form.

7. Known limitations (this build)

8. Path D — viewers / print

Note
EMF is easier than WMF (no placeable-header / aldus-key dance).

Size the target rect from rclFrame at display DPI (§6g). Keep the WMF path for classic pages behind the §2 header check — one page-level branch covers both eras.

8a. A shipped text index and what a viewer can now do (12.0.26)

PageTextIndexClass (LIBSRC ABPREVIEW.INC/.CLW) is the per-page text index the new ReportPreviewClass searches with, and it is GUI-free: a Path A consumer can adopt it as-is instead of walking pages itself. It is the shipped page parser (WMFDocumentParser) driven with a no-op generator that implements both IReportGenerator faces, so every text-bearing control arrives ONCE per control as {text (UTF-16 units on an EMF page, narrow bytes on a WMF page), rectangle in 1/1000 inch paper-absolute, font}; Find(needle, caseSensitive) folds with UPPER on both sides and keeps surrogate pairs whole. Three laws the previewer had to learn, for anyone overlaying a played page:

Known floor: a CHECK caption on a REPORT,UNICODE page is recorded narrowed by the engine (Ω → O); STRING captions are wide. A viewer can now offer what the shipped previewer offers: search with highlighted hits, hit counts per page, marks, print/export marked (export = hand the parser a filtered COPY of the preview queue; the print path prunes).

Compile time (12.0.26): the project define report_unicode=>1 (or PRAGMA('define(report_unicode=>on)') at the top of a module) gives every REPORT the UNICODE attribute as it compiles - the switch for a hand-coded application or a report library that never meets the templates; one report narrows itself with Report{PROP:Unicode} = False before its first PRINT.

Slice 6 (12.0.26): the shipped previewer builds its index on a thread (StartIndexing / TakeIndexed, the ABDOCK START + NOTIFY shape) and copies the runs under a Shift+drag band (CopyRuns: the inverse of the highlight mapping, reading order, wide to the clipboard) - both are VIRTUAL, a viewer of its own can reuse them.

9. Requested feedback and test steps

  1. §5 answers — the wide surface has shipped (IReportGeneratorW, zero-break); what we want to hear now is adoption intent: would your product adopt it, and on which targets first?
  2. Smoke the product on an opted report. Add ,UNICODE to any REPORT (all-ANSI is fine — output should match the classic twin), push pages through the pipeline. Repeat with real wide data (USTRING with mixed script + emoji covers the matrix).
  3. Fixtures available on request. WMF/EMF page-file pairs for the same report (A/B parser ports). Reference EMF arms can be annotated against a product's structure if that helps a port.
  4. Bug reports: page file, extracted record/coordinate/text, expected result, and WMF twin behavior. Code points (U+0041 U+1F600 …) beat mojibake screenshots.

Third-party report tools are why the feature is opt-in per report, not a format flag day. Goal: transition time spent on product-specific code, not rediscovering EMF traps. Detail is either in this note or in the shipped parser source — ask for anything that is neither.