wendashi
# Bug report / fix note: Planner OCR tokens were dropped from the final conditioning prompt **Describe the bug** - **Model I am using (LayoutLM + Diffusion):** - TextDiffuser-2 (t2i full) + TextDiffuser-2 Layout Planner (FastChat-based) - **The problem arises when using:** - [x] the official example scripts: `[inference_textdiffuser2_t2i_full.py](https://github.com/microsoft/unilm/blob/master/textdiffuser-2/inference_textdiffuser2_t2i_full.py)` When running the TextDiffuser-2 inference script(s), the generated prompt debug files (`prompt_{index}_{rank}.txt`) only contained the user prompt and padding tokens, but **did not include** the planner-generated OCR/layout tokens (coordinate tokens + character tokens). As a result, the diffusion model was effectively conditioned **without text layout guidance**, leading to incorrect/weak text placement. --- **To Reproduce** 1. Activate the `textdiffuser-2` environment and run either: - Official script: `python inference_textdiffuser2_t2i_full.py --input_format prompts_txt_file --prompts_txt_file <prompts.txt> --m1_model_path <planner> --output_dir <out>` 2. Open any generated `prompt_{index}_{rank}.txt` (e.g. `prompt_4_-1.txt`). 3. Observe that after the initial `<|startoftext|> ... <|endoftext|>` segment, the rest is mostly repeated `<|endoftext|>` (padding), with **no** `l* t* r* b*` coordinate tokens or `[C]` character tokens. --- Buggy[Left], Fixed [Right] <img width="2464" height="1302" alt="Image" src="https://github.com/user-attachments/assets/1ba5edde-e76b-45cc-b890-14fa27fdec8f" /> **Expected behavior** The prompt debug file should contain: - the caption/user prompt tokens, followed by - planner OCR tokens: coordinate tokens (`l41`, `t24`, `r90`, `b52`, ...) and character tokens (`[S]`, `[h]`, `[o]`, `[o]`, `[t]`, ...) Example (schematic): ``` <|startoftext|>...user prompt...<|endoftext|> l41 t24 r90 b52 [S] [h] [o] [o] [t] <|endoftext|> l49 t48 r82 b73 [Y] [o] [u] [r] <|endoftext|> ... ``` --- **Actual behavior** The prompt debug file contained only the user prompt and padding tokens, e.g.: ``` <|startoftext|>...user prompt...<|endoftext|><|endoftext|><|endoftext|>... ``` --- **Root Cause** - **Files:** - `inference_textdiffuser2_t2i_full.py` - **Location:** around the OCR-token construction loop (line numbers may shift as the files evolve) The script correctly loaded the planner output: ```python current_ocr = ocrs[idx] ``` but then accidentally overwrote it: ```python current_ocr = [] # ❌ clears planner output ``` This made the subsequent loop a no-op: ```python for ocr in current_ocr: ... ``` So `ocr_ids` stayed empty and no layout tokens were appended to the conditioning prompt. --- **Fix** Removed the erroneous overwrite and iterated the planner output as intended. - **Before (buggy):** ```python current_ocr = [] for ocr in current_ocr: ... ``` - **After (fixed):** ```python for ocr in current_ocr: ... ``` --- **Impact** - **Before fix:** prompts lacked OCR/layout tokens → generated images missed/ignored planned text placement. - **After fix:** prompts include OCR/layout tokens → diffusion is conditioned on planned text regions. --- **Platform** - **Platform:** Linux - **Python version:** Python 3.8.20 - **PyTorch version (GPU?):** torch 2.4.1+cu118 - **GPU:** NVIDIA (multi-GPU setup; `--gpu_id` selects device) --- **Suggested commit message** ``` fix(textdiffuser-2): do not clear planner OCR lines before tokenization ```