Log in
Main
Home
News and Events
About TeamWeaver
Team
Recent changes
Tools
Overview
Integrated Search
Woogle
Wiquila
Support
Mailinglists
FAQs
Bugs/Feature requests
Quick links
Source code (SVN)
Tickets (JIRA)
View source
Page
Discussion
View source
History
>>>
From TeamWeaverWiki
for
Repo config.xml
Jump to:
navigation
,
search
The file <code>repo_config.xml</code> in the <code>\teamweaverIS-backend\WEB-INF\conf</code> of your [[Integrated Search/Installation|TeamWeaverIS backend]] allows you to define [[Integrated Search/Supported data sources|data sources]] to be crawled. Users can easily develop an plug in crawlers for other data sources. See our [[Integrated Search/How to write a custom crawler|guide how to write a custom crawler]]. Repositories are crawled by assigning them to a [[crawl_config.xml|crawl]]. == Example <code>repo_config.xml</code> == <pre> <?xml version="1.0" encoding="UTF-8"?> <repo_config> <repositoryInfo> <repoId>1</repoId> <repoName>PushIndexingTestRepo</repoName> <srcType>web</srcType> <srcVersion>1.0</srcVersion> <connectURL>http://www.teamweaver.org/</connectURL> <connectType>http</connectType> <user></user> <pass></pass> <group>all</group> <linkPath></linkPath> <pushIndexAuthKey>test42</pushIndexAuthKey> <pushIndexEnabled>true</pushIndexEnabled> <cacheFulltext>true</cacheFulltext> </repositoryInfo> <repositoryInfo> <repoId>2</repoId> <repoName>files</repoName> <srcType>filesystem</srcType> <srcVersion>1.0</srcVersion> <connectURL>//linde/c$/docs_small</connectURL> <connectType>filesystem</connectType> <linkPath>http://webdav.internal.de/linde/c/docs_small/</linkPath> <user></user> <pass></pass> <group>all</group> </repositoryInfo> </repo_config> </pre> == Documentation of parameters == * <code>repo_config.xml</code> contains <code><repositoryInfo></code> entries for each single repository to be crawled (see example file above) * For specific information about the semantics of parameters in the context of different source types consult the "List of Source Types" below * The parameters inside the <code><repositoryInfo></code> element are as follows: ** <code><repoId>1</repoId></code> - denotes a numerical id for the repository. This needs to be unique within the repositoryInfo elements of the <code>repo_config.xml</code> ** <code><repoName>My Test repository</repoName></code> - a human readable label for the repository ** <code><srcType>web</srcType></code> - denotes the type of repository - see below for a list of allowed keys ** <code><srcVersion>1.0</srcVersion></code> - denotes a version of the repository type - irrelevant for most types ('''OPTIONAL''') ** <code><connectURL>http://www.teamweaver.org/</connectURL></code> - a descriptor of the physical location. The exact form depends on the kind of srcType - e.g. for a "web" ressource, this is a URL, while for a file system, it is a path ** <code><connectType>http</connectType></code> - denotes a conncection mode for srcTypes that allow for a choice - irrelevant for most types ('''OPTIONAL''') ** <code><user>myUser</user></code> - user name, if the srcType requires authentification ('''OPTIONAL''') ** <code><pass>myPass</pass></code> - password ('''OPTIONAL''') ** <code><group>all</group></code> - user group for which the crawled data for this entry should be accessible ('''OPTIONAL''') ** <code><linkPath></linkPath></code> - allows to specify a separate "link path" for repositories, which do not provide "clickable" URLs for the browser. E.g. a network file share might by indexed as <code>\\computer\path\</code> which is typically not clickable in a web result list. Therefore you could provide an alternative link path to the repository (e.g. a WebDAV wrapper) - <code>http://internal.mycompany.de/computer/path/</code> which is then used to refere to results. ('''OPTIONAL''') ** <code><cacheFulltext>false</cacheFulltext></code> - denotes if the backend should cache indexed files in order to make them accessible via the result interface. This is an alternative, if it not possible to expose those systems via the <code><linkPath></code> option. ** <code><pushIndexEnabled>false</pushIndexEnabled></code> - if set to true, this repository can not be actively crawled any more using [[crawl_config.xml]], but will instead push changes to the backend ('''OPTIONAL''') ** <code><pushIndexAuthKey>a_password</pushIndexAuthKey></code> - an arbitrary string which servers for authentification of the push indexing client ('''OPTIONAL''') == List of Source Types == This is the list of allowed <code><srcType></code> attributes for a <code><repositoryInfo></code> entry. The complete authoritative list can be obtained from the [http://svn.polarion.org/repos/community/Teamweaver/Teamweaver/trunk/org.teamweaver.is.backend.crawl/src/main/java/org/teamweaver/is/backend/crawl/aperture/crawler/defaultCrawlers.xml defaultCrawlers.xml] in the SVN. === Web ressources ("web") === * <code><repositoryInfo></code> options ** <code><srcType>web</srcType></code> ** <code><connectURL>http://www.fzi.de/ipe/</connectURL></code> * Behaviour: the web crawler in its current state of implementation starts from the initial page defined by the <code>connectURL</code> and follows all links including resp. starting with <code>connectURL</code> (e.g. <code>http://www.fzi.de/ipe/some_subdirectory/page.html</code> but not <code>http://www.fzi.de/se/page.html</code>) up to a depth of 9 hops. There is currently no way to configure a different behaviour (although this could be implemented fairly easy). === File systems/network shares === * <code><repositoryInfo></code> options ** <code><srcType>filesystem</srcType></code> ** <code><connectType>filesystem</connectType></code> (Example) ** <code><linkPath>http://webdav.internal.de/linde/c/docs_small/</linkPath></code> (Example) === SVN repositories === * <code><repositoryInfo></code> options ** <code><srcType>svn</srcType></code> ** <code><connectURL>http://svn.polarion.org/repos/community/Teamweaver/Teamweaver/trunk/</connectURL></code> (Example) *** <code><connectURL>svn+ssh://anonymous@ontoware.org/svnroot/semweb4j/trunk</connectURL></code> (Alternative Example) ** <code><connectType>http</connectType></code> === CVS repositories === * <code><repositoryInfo></code> options ** <code><srcType>svn</srcType></code> ** <code><connectURL>:pserver:anonymous@aperture.cvs.sourceforge.net:/cvsroot/aperture:aperture</connectURL></code> (Example) ** <code><connectType>pserver</connectType></code> === Atlassian Confluence Wiki === * <code><repositoryInfo></code> options ** <code><srcType>confluence</srcType></code> ** <code><connectURL>http://localhost:8080/confluence/</connectURL></code> (Example) ** <code><connectType>rpc</connectType></code> === Atlassian Jira Issue Tracker === * <code><repositoryInfo></code> options ** <code><srcType>jira</srcType></code> ** <code><connectURL>localhost:8080/</connectURL></code> (Example) ** <code><connectType>rpc</connectType></code> === JSPWiki === * <code><repositoryInfo></code> options ** <code><srcType>jspwiki</srcType></code> ** <code><srcVersion>2.0</srcVersion></code> ** <code><connectURL>http://www.jspwiki.org</connectURL></code> ** <code><connectType>rpc</connectType></code> === Bugzilla Issue Tracker === * <code><repositoryInfo></code> options ** <code><srcType>bugzilla</srcType></code> ** <code><connectURL>http://landfill.bugzilla.org/bugzilla-2.18-branch/</connectURL></code> (Example) ** <code><connectType>rpc</connectType></code> === To be continued.... === * To be continued....
Return to
Repo config.xml
.
Privacy policy
About TeamWeaverWiki
Disclaimers
Powered by MediaWiki
Design by Paul Gu