# Japanese price parsing fixture

Offline synthetic demonstration. No Japanese storefront or proxy was contacted. Run with Python 3.10 or newer (actually verified on Python 3.14.7):

```sh
python3 parse_price.py fixtures.json > replay.json
```

Compare replay.json with results.json: 36 cases, 36 expected classifications, zero failures in the recorded run. Review outcomes are successful abstentions, not failed tests. Each expected field is checked; full actual outputs remain in results.json.

## Input contract

Call `parse_price({"text": "￥１，２９８（税込）", "currency": "JPY", "currency_evidence": "selected_offer.priceCurrency"})`. Select one actual product/variant price field before parsing; supply currency from independent page/API context. `currency_evidence` is a human-supplied reference: the script neither retrieves nor authenticates it. The fixture reference names synthetic context, not a real offer.

Output preserves `original`, a width-normalized string, whole-yen amount and `tax_basis` (`included`, `excluded`, `unspecified`). `parsed` establishes only these input-field checks. It does not establish product identity, availability, a delivered total, source truth or suitability for comparison. Compare like tax bases on the same variant and pricing conditions.

## Deliberately narrow grammar

Accept a single `¥1298`, `1298円`, or `JPY 1298`, optionally followed by a parenthesized Japanese tax label (税込/税込価格/税抜/税抜き/税別), or a space plus that label. Commas must group three digits. No rate or cents multiplication is performed. Whole-yen storage is this parser contract, not a statement about every API or valid Japanese price representation.

Supported width characters are fullwidth ASCII FF01–FF5E, ideographic space U+3000, and fullwidth yen U+FFE5. Other compatibility-changing characters are held before NFKC; this avoids silently folding superscript/circled digits. The backslash U+005C remains a backslash regardless of its font appearance.

Ranges, instalments, points, multiple prices, negative values, decimals, Kanji-number units, a tax label before the price and leading zeros all require review. Some are valid store presentations outside this grammar. Never strip all nondigits from arbitrary page text. Extend the grammar only after inspecting the field and adding independent expected fixtures.

## Sources and date

Prepared 2026-10-06. [Unicode 17 chapter 22](https://www.unicode.org/versions/Unicode17.0.0/core-spec/chapter-22/) documents the yen/yuan glyph ambiguity; [Unicode normalization](https://unicode.org/faq/normalization.html) defines compatibility folding; [Python unicodedata](https://docs.python.org/3/library/unicodedata.html) implements normalization. [Japan NTA 6902](https://www.nta.go.jp/taxes/shiraberu/taxanswer/shohi/6902.htm) gives examples with tax-exclusive and tax-inclusive values together. The executed Python runtime reports Unicode database 16.0.0 in environment.json; these fixtures do not test every Unicode character or version. This demonstration parses labels only; it is not tax advice, a calculator, or a compliance checker.
