Proposal: port moment.js era tokens (`N`…`NNNNN`, `y`, `yo`) to `isoFormat()`, plus two locale format bugs found on the way

#3368 · closed · 3 comments

View on GitHub ↗

kamthorn

## Summary `isoFormat()` follows moment.js tokens, but it lacks moment's era support ([format tokens](https://github.com/moment/momentjs.com/blob/master/docs/moment/04-displaying/01-format.md), [locale `eras`](https://github.com/moment/momentjs.com/blob/master/docs/moment/07-customization/17-eras.md); added in [moment 2.25.0](https://github.com/moment/moment/blob/develop/CHANGELOG.md#2250-see-full-changelog), [#4599](https://github.com/moment/moment/issues/4599)). Without it, there is no way to output a year in a non-Christian era, such as the Thai Buddhist Era (2026 CE = 2569 BE) or Japanese eras, and no way to parse one back. I'd like to contribute this in small PRs. While preparing, I scanned all 824 locales on `master` (`ec211de`) and found two existing bugs. I'd fix those first, because the second one would silently change meaning once a `y` token exists. ## Bug 1: first character after an `L`/`l`/`LT` macro expansion is replaced by the macro letter In `Date::isoFormat()`, after a macro such as `LLLL` is expanded, `$char` still holds `'L'`. If the expanded format starts with literal text (not a token), the first character is output as `L` (or `l`). ```php Carbon::parse('2026-09-30 14:05')->locale('th')->isoFormat('LLLL'); // actual: "Lันพุธที่ 30 กันยายน 2026 เวลา 14:05" // expected: "วันพุธที่ 30 กันยายน 2026 เวลา 14:05" ``` Affected (11 locales, 25 formats, including `calendar.sameElse` where it is `L`): | Locale | Formats | Actual → expected (first chars) | | --- | --- | --- | | `th`, `th_TH` | `LLLL`, `llll` | `Lันพุธ…` → `วันพุธ…` | | `lo`, `lo_LA` | `LLLL`, `llll` | `Lັນພຸດ…` → `ວັນພຸດ…` | | `dz`, `dz_BT` | `L`, `l`, `calendar.sameElse` | `Lསྱི་ལོ…` → `པསྱི་ལོ…` | | `bo_IN` | `LLL`, `lll` | `Lྤྱི་ལོ…` → `སྤྱི་ལོ…` | | `ps`, `ps_AF` | `LLL`, `lll` | `L 2026 د…` → `د 2026 د…` | | `nnh` | `LLL`, `lll` | format starts with `[`, so the escape is lost too: `L30/9/yyyɛ]̌ʼ…` | | `ne_IN` | `L`, `l`, `calendar.sameElse` | caused by Bug 2 (`yy/M/d` starts with a non-token) | Proposed fix: after expanding a macro, re-read the expanded text from its first character, so escapes (`[…]`, `\`) and tokens are handled normally. It's a few lines in `isoFormat()`. ## Bug 2: CLDR pattern letters left in locale data Several `formats` entries use CLDR syntax where the letters mean something else (or nothing) in moment/Carbon tokens: - **`yy` is not a token** and is printed literally, in `L` of 13 files (15 locales): `es_PH`, `fo_DK`, `hr_BA`, `ms_BN`, `ms_SG`, `ne_IN`, `nnh`, `pa_Guru`, `sr_Cyrl_BA`, `sr_Latn_BA`, `ta_MY`, `ta_SG`, `uz_Cyrl`. Example: `es_PH` `L` = `D/M/yy` → `30/9/yy`. - **A lone `d` is the day of the week in moment** but the day of the month in CLDR: `ne_IN` `L` = `yy/M/d` → `yy/9/3`, `ps`/`ps_AF` `L` = `YYYY/M/d` → `2026/9/3`, `seh` `LL`/`LLL`/`LLLL` (`d [de] MMM…` → `3 de Set…`), `nnh` `LLL`/`LLLL`. - **Unescaped words:** `fur` `LL` = `DD di MMMM dal YYYY` → `30 3i setembar 3pm30. 9. 26 2026` (`d`, `a` and `l` in "di"/"dal" are read as tokens). Proposed fix: replace with the equivalent Carbon tokens (`YY`, `D`, `[di]`, …), checked against CLDR, one commit per locale. I'd also add a regression test that scans all locale files for these patterns, so a future `y`/`N` token can't change their meaning silently. ## Feature: era tokens Same semantics and names as moment: | Token | Meaning | `th` | `en` (default) | | --- | --- | --- | --- | | `N`, `NN`, `NNN` | era abbreviation | พ.ศ. | AD | | `NNNN` | era name | พุทธศักราช | Anno Domini | | `NNNNN` | era narrow name | พ.ศ. | AD | | `y` | year of era | 2569 | 2026 | | `yo` | ordinal year of era | 2569 | 2026th | Eras would live in locale files with moment's fields (`since`, `until`, `offset`, `name`, `narrow`, `abbr`), with AD/BC in `en` as the fallback for every locale. Era year = `(year − since.year) × direction + offset`, which is pure arithmetic, needs no ext-intl and doesn't change any date calculation. Parsing (`createFromIsoFormat()` / `createFromLocaleIsoFormat()`) would convert the era year to CE in the input string *before* calling `createFromFormat()`. Parsing as a CE year first and then subtracting would break Feb 29: `createFromFormat('!Y-m-d', '2567-02-29')` gives `2567-03-01`. Planned PRs: (1) Bug 1, (2) Bug 2, (3) era format, (4) era parse. (3) depends on (2). ## Questions 1. **Backward compatibility:** `N` and `y` are currently printed literally when unescaped in user formats. Would adding them as tokens be acceptable in a minor release (moment has had them since 2.25.0, May 2020), or should it wait for a major? 2. Should there be an opt-in way for `th` to use the Buddhist Era in `L`/`LL` (for example a `th@calendar=buddhist`-style variant), or should the defaults stay CE and users write `D MMMM N y` explicitly? I'd keep the defaults unchanged in this work. 3. Is era support in `translatedFormat()` (PHP `date()` letters) of interest? PHP has no free letter for it, so it would need a separate design. I'd leave it out for now. I can open PRs 1 and 2 right away, because they're independent of the answers.

Comments

kamthorn

Some context I should have included: this has come up before, and I'd like to address those earlier answers directly. - #53 (2013) added a Buddhist-era flag to `formatLocalized()`. It was declined as too specific. - #624: "one date library should handle one calendar". - #2954 (2024): `addYears(543)` turned 2024-02-29 into 2567-03-01. A `formatBuddhist()` macro was later added to the docs. How this proposal differs: 1. **No new calendar.** The Thai solar calendar has the same days and months as the Gregorian calendar (since 1941, when the Thai new year moved to January 1). Only the year label differs, like Japanese eras. The date stays a normal Gregorian `DateTime`, and nothing in date arithmetic changes. That keeps the "one library, one calendar" rule. 2. **Not Thai-specific.** It's moment.js's generic era mechanism (tokens `N`…`NNNNN`, `y`, `yo`, plus `eras` in locale data), which `isoFormat()` is modeled on. Thai and Japanese would just be two locales with `eras` data. 3. **It fixes the #2954 pitfall for parsing.** The era year is converted to CE in the input string *before* `createFromFormat()`, so `29 กุมภาพันธ์ พ.ศ. 2567` parses to 2024-02-29. A macro can format, but it can't help `createFromIsoFormat()`.

kylekatarnls

Hi, thanks for the report, can you please open 1 issue per feature/bug so they can be discussed separately, potentially assigned to different releases, and pull-requests to solve them will refer to different issues. Thanks 🙏

kamthorn

Thanks, I've split it into: - #3372 isoFormat(): first character after a macro expansion (fix in #3369) - #3371 CLDR pattern letters in locale formats (fix in #3370) - #3373 era tokens (`N`…`NNNNN`, `y`, `yo`) I'll close this one, so the discussion continues there.