kamthorn
## Summary `isoFormat()` follows moment.js tokens, but it lacks moment's era support ([format tokens](https://github.com/moment/momentjs.com/blob/master/docs/moment/04-displaying/01-format.md), [locale `eras`](https://github.com/moment/momentjs.com/blob/master/docs/moment/07-customization/17-eras.md); added in [moment 2.25.0](https://github.com/moment/moment/blob/develop/CHANGELOG.md#2250-see-full-changelog), [#4599](https://github.com/moment/moment/issues/4599)). Without it, there is no way to output a year in a non-Christian era, such as the Thai Buddhist Era (2026 CE = 2569 BE) or Japanese eras, and no way to parse one back. I'd like to contribute this in small PRs. While preparing, I scanned all 824 locales on `master` (`ec211de`) and found two existing bugs. I'd fix those first, because the second one would silently change meaning once a `y` token exists. ## Bug 1: first character after an `L`/`l`/`LT` macro expansion is replaced by the macro letter In `Date::isoFormat()`, after a macro such as `LLLL` is expanded, `$char` still holds `'L'`. If the expanded format starts with literal text (not a token), the first character is output as `L` (or `l`). ```php Carbon::parse('2026-09-30 14:05')->locale('th')->isoFormat('LLLL'); // actual: "Lันพุธที่ 30 กันยายน 2026 เวลา 14:05" // expected: "วันพุธที่ 30 กันยายน 2026 เวลา 14:05" ``` Affected (11 locales, 25 formats, including `calendar.sameElse` where it is `L`): | Locale | Formats | Actual → expected (first chars) | | --- | --- | --- | | `th`, `th_TH` | `LLLL`, `llll` | `Lันพุธ…` → `วันพุธ…` | | `lo`, `lo_LA` | `LLLL`, `llll` | `Lັນພຸດ…` → `ວັນພຸດ…` | | `dz`, `dz_BT` | `L`, `l`, `calendar.sameElse` | `Lསྱི་ལོ…` → `པསྱི་ལོ…` | | `bo_IN` | `LLL`, `lll` | `Lྤྱི་ལོ…` → `སྤྱི་ལོ…` | | `ps`, `ps_AF` | `LLL`, `lll` | `L 2026 د…` → `د 2026 د…` | | `nnh` | `LLL`, `lll` | format starts with `[`, so the escape is lost too: `L30/9/yyyɛ]̌ʼ…` | | `ne_IN` | `L`, `l`, `calendar.sameElse` | caused by Bug 2 (`yy/M/d` starts with a non-token) | Proposed fix: after expanding a macro, re-read the expanded text from its first character, so escapes (`[…]`, `\`) and tokens are handled normally. It's a few lines in `isoFormat()`. ## Bug 2: CLDR pattern letters left in locale data Several `formats` entries use CLDR syntax where the letters mean something else (or nothing) in moment/Carbon tokens: - **`yy` is not a token** and is printed literally, in `L` of 13 files (15 locales): `es_PH`, `fo_DK`, `hr_BA`, `ms_BN`, `ms_SG`, `ne_IN`, `nnh`, `pa_Guru`, `sr_Cyrl_BA`, `sr_Latn_BA`, `ta_MY`, `ta_SG`, `uz_Cyrl`. Example: `es_PH` `L` = `D/M/yy` → `30/9/yy`. - **A lone `d` is the day of the week in moment** but the day of the month in CLDR: `ne_IN` `L` = `yy/M/d` → `yy/9/3`, `ps`/`ps_AF` `L` = `YYYY/M/d` → `2026/9/3`, `seh` `LL`/`LLL`/`LLLL` (`d [de] MMM…` → `3 de Set…`), `nnh` `LLL`/`LLLL`. - **Unescaped words:** `fur` `LL` = `DD di MMMM dal YYYY` → `30 3i setembar 3pm30. 9. 26 2026` (`d`, `a` and `l` in "di"/"dal" are read as tokens). Proposed fix: replace with the equivalent Carbon tokens (`YY`, `D`, `[di]`, …), checked against CLDR, one commit per locale. I'd also add a regression test that scans all locale files for these patterns, so a future `y`/`N` token can't change their meaning silently. ## Feature: era tokens Same semantics and names as moment: | Token | Meaning | `th` | `en` (default) | | --- | --- | --- | --- | | `N`, `NN`, `NNN` | era abbreviation | พ.ศ. | AD | | `NNNN` | era name | พุทธศักราช | Anno Domini | | `NNNNN` | era narrow name | พ.ศ. | AD | | `y` | year of era | 2569 | 2026 | | `yo` | ordinal year of era | 2569 | 2026th | Eras would live in locale files with moment's fields (`since`, `until`, `offset`, `name`, `narrow`, `abbr`), with AD/BC in `en` as the fallback for every locale. Era year = `(year − since.year) × direction + offset`, which is pure arithmetic, needs no ext-intl and doesn't change any date calculation. Parsing (`createFromIsoFormat()` / `createFromLocaleIsoFormat()`) would convert the era year to CE in the input string *before* calling `createFromFormat()`. Parsing as a CE year first and then subtracting would break Feb 29: `createFromFormat('!Y-m-d', '2567-02-29')` gives `2567-03-01`. Planned PRs: (1) Bug 1, (2) Bug 2, (3) era format, (4) era parse. (3) depends on (2). ## Questions 1. **Backward compatibility:** `N` and `y` are currently printed literally when unescaped in user formats. Would adding them as tokens be acceptable in a minor release (moment has had them since 2.25.0, May 2020), or should it wait for a major? 2. Should there be an opt-in way for `th` to use the Buddhist Era in `L`/`LL` (for example a `th@calendar=buddhist`-style variant), or should the defaults stay CE and users write `D MMMM N y` explicitly? I'd keep the defaults unchanged in this work. 3. Is era support in `translatedFormat()` (PHP `date()` letters) of interest? PHP has no free letter for it, so it would need a separate design. I'd leave it out for now. I can open PRs 1 and 2 right away, because they're independent of the answers.