Pulpie is a family of encoder models that clean the web at scale. Match state-of-the-art extraction quality while cutting costs from $159k to $7.9k per billion pages.
Start Free TrialCost to clean 1 billion webpages:
$7,900 with Pulpie vs $159,000 with Dripper
One forward pass over full HTML, not token-by-token generation. Compute-bound efficiency instead of memory-bound decoder bottlenecks.
Matches leading extractors on boilerplate removal, content preservation, and structural accuracy across 10k+ real-world sites.
Choose trade-offs: speed for bulk crawls, accuracy for curated content, or balanced extraction for general-purpose pipelines.
Output cleaned HTML for downstream processing or Markdown for direct LLM consumption. Preserve structure or simplify.
Run inference on CPU or GPU. Batch process millions of pages per day without expensive server clustering.
Self-host or use our API. No vendor lock-in, no rate limits, full control over data privacy and cost.
Pulpie models use a transformers-based encoder that jointly learns to classify HTML blocks as content, boilerplate, or structural. Each block receives a binary label in a single forward pass, avoiding the cumulative latency of decoder-based approaches. By parallelizing block classification across the full DOM tree, inference scales linearly with page size and runs efficiently on modest hardware. Training leverages human-curated extraction datasets and adversarial examples from real-world edge cases.
| Metric | Pulpie | Dripper | Savings |
|---|---|---|---|
| 1M pages | $7.90 | $159 | 95% cheaper |
| 1B pages | $7,900 | $159,000 | $151,100 saved |
| Quality Score | 98% | 98% | Equivalent |
| Inference Latency | 120ms / page | 280ms / page | 2.3x faster |
Crawl millions of URLs daily. Pulpie's encoder runs inference 20x cheaper than traditional NLP pipelines. Perfect for search quality teams, SERP monitoring, and competitive intelligence.
Clean web text at scale for training and fine-tuning. Remove ads, navigation, and boilerplate without expensive human review. Ship curated datasets faster and cheaper.
Strip formatting from partner websites. Pulpie preserves readability while removing visual cruft, enabling clean redistribution and licensing workflows.
Index internal docs, knowledge bases, and wikis. Extract clean text for retrieval-augmented generation. Run inference on-prem for privacy-sensitive organizations.
Pulpie uses learned models instead of hand-coded heuristics. It generalizes to messy real-world HTML that rule-based tools struggle with. Costs scale better for massive crawls.
Yes. Download the open-weight models and run inference locally. We provide Docker images and Python/Node.js SDKs. No cloud dependency required.
Pulpie matches or exceeds SOTA on common benchmarks (Common Crawl, news sites, academic papers). Our models are trained on human-curated examples from diverse domains.
Input: Raw HTML, gzipped HTML, HTML5 files. Output: Clean HTML, Markdown, structured JSON with metadata (title, author, publish date).
Yes. Enterprise customers processing 10B+ pages/month get per-page rates as low as $0.000001. Contact sales for a quote.
Start with our free tier. Extract up to 1M pages/month. Upgrade to production pricing when you're ready to scale.
Get API KeyDocumentation available at docs.pulpie.dev
Register interest
This is not a purchase and there is no card field. It puts your address, this product, and whatever you write below in front of a person, and you get a written answer about what finishing it, or handing it over for you to run yourself, would actually take.