Log in
Main
Home
News and Events
About TeamWeaver
Team
Recent changes
Tools
Overview
Integrated Search
Woogle
Wiquila
Support
Mailinglists
FAQs
Bugs/Feature requests
Quick links
Source code (SVN)
Tickets (JIRA)
View source
Page
Discussion
View source
History
>>>
From TeamWeaverWiki
for
Integrated Search/How to write a custom crawler
Jump to:
navigation
,
search
This short guide describes how to write crawlers to make additional data sources available for searching with TeamWeaver [[Integrated Search]]. A list of existing crawlers can be found at: [[Integrated_Search/Supported_data_sources|supported data sources]] resp. [[repo_config.xml]]. There are two general strategies for crawling data sources with TeamWeaverIS: * "Pull"-Crawlers are included in TeamWeaverIS and allow to crawl/extract data by accessing external sources/system. Pull-crawlers are fairly easy to realize, but have the disadvantage, that the index might not be accurate, since TeamWeaverIS does only learn about data changes, when a crawl is executed. * "Push"-Crawlers are implemented on the side of the client system which includes the data. They proactively update the index and thus have a direct connection to the TeamWeaverIS backend. Accordingly, push-crawlers are more difficult to implement, but can reflect changes in the data more rapidly in the index. == Writing a "Pull"-Crawler == At the basic level, creating a pull-crawler requires to implement two Java classes and changing two XML-files - a ''Crawler'' and a ''Processor'' each. Crawlers are classes which access a data source and extract/create single data items. Processors act upon these items to prepare them for feeding into the index. === Crawler === * You need to create a subclass of [http://aperture.sourceforge.net/doc/javadoc/org/semanticdesktop/aperture/crawler/base/CrawlerBase.html CrawlerBase] which basically means to implement a method <code>crawlObject</code>. See our [http://svn.polarion.org/repos/community/Teamweaver/Teamweaver/trunk/org.teamweaver.is.backend.crawl/src/main/java/org/teamweaver/is/backend/crawl/aperture/crawler/JiraCrawler.java JIRACrawler] for an example. * Afterwards, register your new crawler in [http://svn.polarion.org/repos/community/Teamweaver/Teamweaver/trunk/org.teamweaver.is.backend.crawl/src/main/java/org/teamweaver/is/backend/crawl/aperture/crawler/defaultCrawlers.xml defaultCrawlers.xml]. You need to define a unique <code><repoType></code> key, which corresponds to the <code><srcType></code> in [[repo_config.xml]]. === Processor === * You need to create a subclass of [http://svn.polarion.org/repos/community/Teamweaver/Teamweaver/trunk/org.teamweaver.is.backend.crawl/src/main/java/org/teamweaver/is/backend/crawl/processor/impl/ProcessorBase.java ProcessorBase]. See our [http://svn.polarion.org/repos/community/Teamweaver/Teamweaver/trunk/org.teamweaver.is.backend.crawl/src/main/java/org/teamweaver/is/backend/crawl/processor/impl/JiraProcessor.java JIRAProcessor] as an example. * Afterwards, register your new processor in [http://svn.polarion.org/repos/community/Teamweaver/Teamweaver/trunk/org.teamweaver.is.backend.crawl/src/main/java/org/teamweaver/is/backend/crawl/processor/impl/defaults.xml defaults.xml] by wiring it with the corresponding crawler. === Advanced topics === * tbd. (generic JDBC crawler, Plain index vs. metadata) == Writing a "Push"-Crawler == * tbd. * Push-Crawlers have to call the TeamWeaver backend's [http://svn.polarion.org/repos/community/Teamweaver/Teamweaver/trunk/org.teamweaver.is.api/src/main/java/org/teamweaver/is/api/IndexService.java IndexService] * You might want to look at our [[Woogle]] code for an example implementation of a push-crawler
Return to
Integrated Search/How to write a custom crawler
.
Privacy policy
About TeamWeaverWiki
Disclaimers
Powered by MediaWiki
Design by Paul Gu