newspaper4k

Newspaper4k is a Python library designed for extracting, processing, and analyzing news articles from websites. It is a continuation and active fork of the original newspaper3k library, which had stopped receiving updates, with the goal of keeping the ecosystem maintained while adding improvements and bug fixes. It provides developers with tools to automatically download web pages, extract the main article content, and collect associated metadata such as titles, authors, images, and publication dates. Newspaper4k also includes natural language processing capabilities that can generate summaries and identify keywords from extracted article text. Newspaper4k supports both single-article extraction and full news site processing, allowing users to build sources representing entire publications and iterate through their articles. It maintains compatibility with the original project so that existing code written for newspaper3k can continue working with minimal changes.

Features

Extracts full article text, titles, authors, and publication dates
Retrieves images, videos, and other metadata from news pages
Supports keyword extraction and article summarization using NLP
Processes individual articles or entire news websites as sources
Provides a Python API and command-line interface for scraping tasks
Maintains compatibility with the original newspaper3k library

Project Samples

Project Activity

See All Activity >

License

MIT License

Follow newspaper4k

newspaper4k Web Site

Other Useful Business Software

Stop Storing Third-Party Tokens in Your Database

Auth0 Token Vault handles secure token storage, exchange, and refresh for external providers so you don't have to build it yourself.

Rolling your own OAuth token storage can be a security liability. Token Vault securely stores access and refresh tokens from federated providers and handles exchange and renewal automatically. Connected accounts, refresh exchange, and privileged worker flows included.

Try Auth0 for Free

Rate This Project

User Reviews

Be the first to post a review of newspaper4k!

Additional Project Details

Programming Language

Python, Unix Shell

Related Categories

Unix Shell Web Scrapers, Python Web Scrapers

Registered

13 hours ago

Similar Business Software

Apify

Apify is a full-stack web scraping and automation platform helping anyone get value from the web. At its core is Apify Store, a marketplace with over 10,000 Actors where developers build, publish, and monetize automation tools. Actors are serverless cloud programs that extract data, automate...

See Software
Oxylabs

Oxylabs is a market leader in web intelligence with enterprise-grade, ethical, and compliant solutions. Its proxy infrastructure spans one of the largest global networks, offering residential, ISP, mobile, datacenter, & dedicated datacenter proxies, along with Web Unblocker – an AI-driven...

See Software
NetNut

Get ready to experience unmatched control and insights with our user-friendly dashboard tailored to your needs. Monitor and adjust your proxies with just a few clicks. Track your usage and performance with detailed statistics. Our team is devoted to providing customers with proxy solutions...

See Software
PYPROXY

Market-leading proxy solution provides tens of millions of IP resources. Commercial residential and ISP proxy network includes 90M+ IPs around the world. Exclusive high-performance server requests access to real residential addresses. Abundant bandwidth support business demands. Real-time speed...

See Software
Price2Spy

Price2Spy makes automatic price adjustments easy to perform saving your most valuable resource - time, allowing your pricing team to focus on strategic planning and management. Since 2010, we have provided pricing intelligence for retailers and brands in 40+ countries, helping them smoothly...

See Software
MangoProxy

MangoProxy is a professional residential proxy service designed for developers, web scrapers, and traffic arbitrage specialists. Key Features: • 90M+ residential IP addresses from 200+ countries • API integration for Python, JavaScript, Go, and other languages • Automatic IP rotation to...

See Software

Report inappropriate content

newspaper4k

Python library for scraping and analyzing online news articles easily

Get an email when there's a new version of newspaper4k

Features

Project Samples

Project Activity

Categories

License

Follow newspaper4k

User Reviews

Additional Project Details

Programming Language

Related Categories

Registered