Start building on Google Cloud with $300 in free credits. No commitment, no credit card required until you're ready to scale.
Launch your next project with $300 in free Google Cloud credits—no strings attached. Test, build, and deploy without risk. Use your credits across the entire Google Cloud platform to find what works best for your needs. After your credits are used, continue with always-free tier services. Only pay when you're ready to scale. Sign up in minutes and start exploring.
Start Free Trial
Build Agents and Models on One Platform
Everything you need to build production-ready agents and models. Access 200+ Google and third-party AI models and tools.
Gemini Enterprise Agent Platform is Google Cloud's comprehensive platform for developers to build, scale, govern, and optimize agents and models. Choose from Google's most advanced models and third-party models like Anthropic's Claude Model Family.
Easy Spider is a distributed Perl Web Crawler Project from 2006
Easy Spider is a distributed Perl Web Crawler Project from 2006. It features code from crawling webpages, distributing it to a server and generating xml files from it. The client site can be any computer (Windows or Linux) and the Server stores all data.
Websites that use EasySpider Crawling for Article Writing...
Regain is a Java search engine based on Jakarta Lucene. It provides indexing and searching files for plenty of formats (HTML,XML,doc(x),xls(x),ppt(x),oo,PDF,RTF,mp3,mp4,Java). A TagLibrary eases integrating search results in your JSP based web page.
IDRA (InDexing and Retrieving Automatically) is a tool which allows indexing a wide range of text (TXT, DOC, PDF) and image annotations files (XML), query-based searching, visualizing an index, saving it for re-usability, evaluation, etc.
Group file share with advanced text parsing capability for easy search
Originally created as a church resource sharing system, phpShare&Search allows users to create accounts, share documents, search documents, and like or report documents.
phpShare&Search's power comes from its advanced document parser which extracts text from .PDF, .TXT, .DOC, and .DOCX files and its community features of liking resources and reporting them as inappropriate or SPAM. Users also subscribe to weekly updates of new content.
User's may choose to download and host/install/configure/modify/manage this code themselves, or contract the code writer to do these functions for them. Contact me for a reasonable quote. eedrew <at> users <dot> sourceforge <dot> net
To support future revisions and/or contribute based on the value you found from this code, checkout the External Link drop-down in the menu.
...
DocInfoRetriever is a Web_based document full-text search engine based on lucene. It allows you to search the contents and metadata of documents . Supported document formats, likes doc, xls, pdf, odt, jpg...etc.,and torrent files.
Mesin pencari berkas doc, docx, ppt, pptx dan pdf (Open Source).
Created by : X-Cisadane (Dwi). Greetz to : XCode, Dunia Santai, Depok Cyber, Borneo Crew, Muslim Hackers, Hacker Cisadane, UG-HotZone 567.
...This php class allow websites developpers to use a powerfull cache system to increase server performance. This is an early version. Futures versions will contain an admin tool. Note that the french doc included is not over yet.
An intranet based document indexing/search facility. Creates an index of MS Office documents (.doc, .xls, .ppt) plain text and .PDF files found in the UNC path passed to the script.Results are store in MySQL database with PHP frontend.
KSearch website search engine, written in Perl, is fully customizable with unlimited page search. Can use DBM or flat-file database. Search results output produce XHTML 1.0 Strict doc types making HTML and CSS easily match your existing website.
Letter Doc will be a system to be able to view and manage a variety of document of different file formats. It is going to be web-based, and built with PHP and MySQL initially.
CitemaPP is a Google Sitemap generator written in C++. Instead of crawling the html-doc directory on your server, CitemaPP crawls the content of your server via http protocol.
Atlantide is a PHP CMS for managing Role Playing Games related stuff.
It aims to be an good-structured archive with a good search engine (with in-doc fulltext search) and embeddable in major portal CMS as various flavors of Nuke.
This projects implements a complete entreprise solution based on lucene. It's a smart engine implemented to index numerous files formats (pdf, ps, xls, doc, ppt, ). The engine can index file systems (filtering), databases, mailing folders, web sites and
E-Xoops Digger is an advanced search engine for Xoops and E-Xoops. Features are content indexing (like pdf, doc, xls), display results by rank, limit to a certain module, results with text spnippet highlighted, fuzzy search, exact search....
yaDMS, which stands for yet another Document Management System, is a php based DMS, with many Features like Clipboard, Mail2DMS, DMS2Mail, Zip&Download, Copy, Move, Multiuser, Fulltextsearch (doc,pdf,rtf,txt,mp3, external tools needed).
Ce projet est une modification du composant DocMan 1.4.0 pour Joomla 1.5.9, qui permet d'ajouter une recherche plein texte sur les fichiers de type: PPT, PDF, DOC, TXT, HTML, PS.