MediaCrawler is an open-source tool for collecting public content from major Chinese social media platforms. It supports Xiaohongshu, Douyin, Kuaishou, Bilibili, Weibo, Tieba, Zhihu, and other services. The project uses Playwright to automate browser login and preserve authenticated sessions. It obtains required request signatures from the active browser context, avoiding complex JavaScript reverse engineering. Users can search by keyword, crawl specific posts, collect nested comments, and retrieve creator profiles. It also supports cached login states, proxy pools, and comment word-cloud generation. The repository is intended for learning and research, with clear restrictions against illegal, large-scale, or commercial scraping.

Features

  • Multi-platform social media crawling
  • Keyword and post-ID collection
  • Nested comment extraction
  • Creator profile retrieval
  • Cached browser login sessions
  • Proxy and word-cloud support

Project Samples

Project Activity

See All Activity >

Categories

Multimedia

Follow MediaCrawler

MediaCrawler Web Site

Other Useful Business Software
Build Agents and Models on One Platform Icon
Build Agents and Models on One Platform

Everything you need to build production-ready agents and models. Access 200+ Google and third-party AI models and tools.

Gemini Enterprise Agent Platform is Google Cloud's comprehensive platform for developers to build, scale, govern, and optimize agents and models. Choose from Google's most advanced models and third-party models like Anthropic's Claude Model Family.
Try It Free
Rate This Project
Login To Rate This Project

User Reviews

Be the first to post a review of MediaCrawler!

Additional Project Details

Operating Systems

Linux, Mac, Windows

Programming Language

Python

Related Categories

Python Multimedia Software

Registered

2 days ago