This project implements an OpenEnv-compatible reinforcement learning environment that simulates a real-world document formatting task:
Convert structured content into HTML such that the rendered PDF matches a reference document in content, structure, and layout.
This environment evaluates practical agent capabilities such as:
- Structured generation
- Formatting correctness
- Layout reasoning
- Iterative improvement
Most RL environments are toy problems. This environment focuses on:
- Real-world document workflows (resumes, reports, structured docs)
- HTML โ PDF rendering correctness
- Semantic + structural + layout evaluation
- Dense feedback for learning agents
DocumentAction:
html_code: strAgent outputs HTML representing the document.
DocumentObservation:
task_description
render_success
render_log
extracted_text
score
pdf_score_breakdown
html_score_breakdown
best_score
attempt_numberAgent receives:
- Rendering feedback
- Extracted PDF content
- Score + breakdown
- Progress signal
DocumentState:
episode_id
task_description
step_count
best_score
task_id
difficulty
referenceTracks task + progress across steps.
Each step() executes:
- Sanity checks on generated HTML
- Converts HTML to PDF
- Text extraction
- Layout extraction (bounding boxes, fonts, spacing)
- Keyword matching (exact + fuzzy)
- Section detection
- Section ordering
- Anti-spam penalty
- Hierarchy validation (heading โ paragraph)
- Alignment consistency
- Spacing correctness
- Content density checks
- Structural validation (partial)
Dense reward in range [0.0, 1.0]
reward =
0.3 * html_score +
0.4 * text_score +
0.3 * layout_score- Continuous reward (not sparse)
- Encourages partial correctness
- Penalizes spam / bad formatting
- Supports iterative refinement
3 task levels:
- Keyword + section presence
- Structure + ordering
- Full layout + formatting quality
Each task includes:
- Prompt
- Reference content
- Expected keywords
- Expected sections
reset() โ new task
step(action) โ evaluate HTML
state() โ current stateEpisode ends when:
- Max steps reached OR
- Score threshold achieved
Deterministic and reproducible.
- Keyword coverage
- Section presence
- Ordering correctness
- Spam penalty
- Hierarchy correctness
- Alignment
- Spacing
- Density
Includes inference.py:
- Uses OpenAI-compatible API
- Outputs:
[START]
[STEP]
[END]
- Computes:
- Step rewards
- Final normalized score (0โ1)
uvicorn server.app:app --host 0.0.0.0 --port 8000docker build -t document-env .
docker run -p 8000:8000 document-envenv.reset()
env.step(DocumentAction(html_code="<html>...</html>"))- Hugging Face Spaces compatible
- WebSocket-based sessions
- Supports concurrent environments
| Criteria | Status |
|---|---|
| Real-world utility | โ |
| 3 tasks | โ |
| Deterministic graders | โ |
| Dense reward | โ |
| OpenEnv spec | โ |
| Baseline script | โ |
| Docker | โ |
| HF deployment | ๐ง |
This environment evaluates:
Structured document generation with layout correctness
A capability where current LLMs are weak.
- Improve HTML grading
- Add multi-page layout support
- Expand task diversity
- Add stricter layout constraints
- Real-world RL environment
- Dense, meaningful reward
- Deterministic evaluation
- OpenEnv compliant
- Scalable deployment ready
Meta x PyTorch OpenEnv Hackathon Submission ๐