HalfSweet/ChopText

ChopText is a smart text segmentation tool designed to split long text into smaller chunks based on a maximum character limit

★ 0Forks 0GitHub ↗Compare

README

ChopText

ChopText is a smart text segmentation tool designed to split long text into smaller chunks based on a maximum character limit, while preserving sentence structure and paragraph integrity. Built with modern web technologies for optimal performance and user experience.

Features

  • Smart Splitting: Intelligent segmentation that respects punctuation (.?!;), commas, and spaces.
  • Customizable Length: Set your preferred maximum characters per chunk (N).
  • Paragraph Handling: Choose to split by single newlines or standard double-newline paragraphs.
  • Internationalization: Fully localized for English and Chinese (Simplified).
  • Dark Mode: Supports system, light, and dark themes.
  • Privacy Focused: All processing happens locally in your browser.
  • Export: Copy to clipboard or download as .txt.

Tech Stack

  • Framework: React + TypeScript + Vite
  • UI Library: shadcn/ui + Tailwind CSS
  • State/i18n: React Hooks + react-i18next
  • Testing: Vitest
  • Tooling: ESLint, Prettier, PostCSS

Getting Started

Prerequisites

  • Node.js 18+
  • pnpm

Installation

  1. Clone the repository:

    git clone https://github.com/halfsweet/ChopText.git
    cd ChopText
  2. Install dependencies:

    pnpm install
  3. Start development server:

    pnpm dev
  4. Build for production:

    pnpm build
  5. Run tests:

    pnpm test

Deployment

This project is configured to automatically deploy to GitHub Pages via GitHub Actions.

  1. Go to your repository Settings -> Pages.
  2. Select GitHub Actions as the source.
  3. Push your code to the main branch.
  4. The workflow will automatically build and deploy the site.

Note: The vite.config.ts is configured with base: "/ChopText/". If your repository name is different, please update the base path in vite.config.ts.

Algorithm Details

The ChopText algorithm prioritizes cutting text at natural boundaries:

  1. Strong punctuation (running sentence endings like ., ?, !, 。).
  2. Weak punctuation (clauses like ,, ,).
  3. Spaces (for Western languages).
  4. Forced hard cut (if no boundaries found within the limit).

It normalizes CJK and mixed text handling by checking for standard Unicode punctuation marks.

License

Apache License 2.0

Contributors

HalfSweet

Issues