MediaCrawler is an open-source tool for collecting public content from major Chinese social media platforms. It supports Xiaohongshu, Douyin, Kuaishou, Bilibili, Weibo, Tieba, Zhihu, and other services. The project uses Playwright to automate browser login and preserve authenticated sessions. It obtains required request signatures from the active browser context, avoiding complex JavaScript reverse engineering. Users can search by keyword, crawl specific posts, collect nested comments, and retrieve creator profiles. It also supports cached login states, proxy pools, and comment word-cloud generation. The repository is intended for learning and research, with clear restrictions against illegal, large-scale, or commercial scraping.

Features

  • Multi-platform social media crawling
  • Keyword and post-ID collection
  • Nested comment extraction
  • Creator profile retrieval
  • Cached browser login sessions
  • Proxy and word-cloud support

Project Samples

Project Activity

See All Activity >

Categories

Multimedia

Follow MediaCrawler

MediaCrawler Web Site

Other Useful Business Software
$300 Free Credits for Your Google Cloud Projects Icon
$300 Free Credits for Your Google Cloud Projects

Start building on Google Cloud with $300 in free credits. No commitment, no credit card required until you're ready to scale.

Launch your next project with $300 in free Google Cloud credits—no strings attached. Test, build, and deploy without risk. Use your credits across the entire Google Cloud platform to find what works best for your needs. After your credits are used, continue with always-free tier services. Only pay when you're ready to scale. Sign up in minutes and start exploring.
Start Free Trial
Rate This Project
Login To Rate This Project

User Reviews

Be the first to post a review of MediaCrawler!

Additional Project Details

Operating Systems

Linux, Mac, Windows

Programming Language

Python

Related Categories

Python Multimedia Software

Registered

2 days ago