iamsyg/doc-generation

โ˜… 0Forks 0PythonGitHub โ†—Compare

README

๐Ÿ“„ Document-to-PDF RL Environment (OpenEnv)

๐Ÿš€ Overview

This project implements an OpenEnv-compatible reinforcement learning environment that simulates a real-world document formatting task:

Convert structured content into HTML such that the rendered PDF matches a reference document in content, structure, and layout.

This environment evaluates practical agent capabilities such as:

  • Structured generation
  • Formatting correctness
  • Layout reasoning
  • Iterative improvement

๐ŸŽฏ Motivation

Most RL environments are toy problems. This environment focuses on:

  • Real-world document workflows (resumes, reports, structured docs)
  • HTML โ†’ PDF rendering correctness
  • Semantic + structural + layout evaluation
  • Dense feedback for learning agents

๐Ÿ—๏ธ Environment Design

๐Ÿ”น Action Space

DocumentAction:
    html_code: str

Agent outputs HTML representing the document.


๐Ÿ”น Observation Space

DocumentObservation:
    task_description
    render_success
    render_log
    extracted_text
    score
    pdf_score_breakdown
    html_score_breakdown
    best_score
    attempt_number

Agent receives:

  • Rendering feedback
  • Extracted PDF content
  • Score + breakdown
  • Progress signal

๐Ÿ”น State

DocumentState:
    episode_id
    task_description
    step_count
    best_score
    task_id
    difficulty
    reference

Tracks task + progress across steps.


๐Ÿ”„ Environment Workflow

Each step() executes:

1. HTML Validation

  • Sanity checks on generated HTML

2. HTML โ†’ PDF Rendering

  • Converts HTML to PDF

3. Extraction

  • Text extraction
  • Layout extraction (bounding boxes, fonts, spacing)

4. Grading

๐Ÿ“Œ Text Grader

  • Keyword matching (exact + fuzzy)
  • Section detection
  • Section ordering
  • Anti-spam penalty

๐Ÿ“Œ Layout Grader

  • Hierarchy validation (heading โ†’ paragraph)
  • Alignment consistency
  • Spacing correctness
  • Content density checks

๐Ÿ“Œ HTML Grader

  • Structural validation (partial)

๐ŸŽฏ Reward Function

Dense reward in range [0.0, 1.0]

reward =
    0.3 * html_score +
    0.4 * text_score +
    0.3 * layout_score

Properties

  • Continuous reward (not sparse)
  • Encourages partial correctness
  • Penalizes spam / bad formatting
  • Supports iterative refinement

๐Ÿง  Tasks

3 task levels:

๐ŸŸข Easy

  • Keyword + section presence

๐ŸŸก Medium

  • Structure + ordering

๐Ÿ”ด Hard

  • Full layout + formatting quality

Each task includes:

  • Prompt
  • Reference content
  • Expected keywords
  • Expected sections

๐Ÿ” Episode Lifecycle

reset() โ†’ new task
step(action) โ†’ evaluate HTML
state() โ†’ current state

โœ… Termination

Episode ends when:

  • Max steps reached OR
  • Score threshold achieved

๐Ÿ“Š Grading System

Deterministic and reproducible.

Text Score

  • Keyword coverage
  • Section presence
  • Ordering correctness
  • Spam penalty

Layout Score

  • Hierarchy correctness
  • Alignment
  • Spacing
  • Density

๐Ÿงช Baseline Inference

Includes inference.py:

  • Uses OpenAI-compatible API
  • Outputs:
[START]
[STEP]
[END]
  • Computes:
    • Step rewards
    • Final normalized score (0โ€“1)

โš™๏ธ Setup & Usage

Run Locally

uvicorn server.app:app --host 0.0.0.0 --port 8000

Docker

docker build -t document-env .
docker run -p 8000:8000 document-env

Client Usage

env.reset()
env.step(DocumentAction(html_code="<html>...</html>"))

โ˜๏ธ Deployment

  • Hugging Face Spaces compatible
  • WebSocket-based sessions
  • Supports concurrent environments

๐Ÿ“ˆ Hackathon Alignment

Criteria Status
Real-world utility โœ…
3 tasks โœ…
Deterministic graders โœ…
Dense reward โœ…
OpenEnv spec โœ…
Baseline script โœ…
Docker โœ…
HF deployment ๐Ÿšง

๐Ÿง  Key Insight

This environment evaluates:

Structured document generation with layout correctness

A capability where current LLMs are weak.


๐Ÿšง Future Work

  • Improve HTML grading
  • Add multi-page layout support
  • Expand task diversity
  • Add stricter layout constraints

๐Ÿ Summary

  • Real-world RL environment
  • Dense, meaningful reward
  • Deterministic evaluation
  • OpenEnv compliant
  • Scalable deployment ready

Meta x PyTorch OpenEnv Hackathon Submission ๐Ÿš€

Contributors

iamsygAkashdeep-ofc

Issues