scrape-it is a Node.js library for converting web pages and HTML documents into structured JavaScript objects. Developers describe the desired output through a declarative schema based on CSS selectors. Fields can extract text, raw HTML, element attributes, repeated items, and nested lists. Conversion functions can transform captured values into dates, numbers, or other application-specific types. The library supports promises, async and await workflows, and callback-based usage. It can request ordinary pages directly or parse HTML obtained from local files and headless browsers. Scrape It does not execute client-side JavaScript itself, so dynamic sites may require an API endpoint or a separate browser automation tool.

Features

  • Declarative CSS selector schemas
  • Text, HTML, and attribute extraction
  • Nested list scraping
  • Custom value conversion functions
  • Remote and local HTML processing
  • Promise and async workflow support

Project Samples

Project Activity

See All Activity >

Categories

Web Scrapers

License

MIT License

Follow scrape-it

scrape-it Web Site

Other Useful Business Software
Build Agents and Models on One Platform Icon
Build Agents and Models on One Platform

Everything you need to build production-ready agents and models. Access 200+ Google and third-party AI models and tools.

Gemini Enterprise Agent Platform is Google Cloud's comprehensive platform for developers to build, scale, govern, and optimize agents and models. Choose from Google's most advanced models and third-party models like Anthropic's Claude Model Family.
Try It Free
Rate This Project
Login To Rate This Project

User Reviews

Be the first to post a review of scrape-it!

Additional Project Details

Operating Systems

Linux

Programming Language

JavaScript

Related Categories

JavaScript Web Scrapers

Registered

5 days ago