Skip to main content

Knowledge bases

A knowledge base is a named collection of sources that Kumkuat crawls on a schedule and keeps searchable. Link one to a persona and its most relevant recent passages are retrieved into that persona's prompt.

Source types

TypeWhat Kumkuat fetches
URL / newsroomPages under a site, crawled by Kumkuat's own crawler with a hosted crawl service as fallback for sites that block plain crawlers
RSSNew items as they are published
SearchA saved web search, rerun on schedule to find recent coverage of a person, company or topic
SocialPublic posts and profile data from X, LinkedIn, Bluesky and Truth Social, via public APIs and a dataset partner
DriveA connected Google Drive or SharePoint library, read-only, with change detection (see below)

Every source shows its last crawl, document count and status. Discover on a newsroom knowledge base queues a search for the newsroom's feeds and section pages.

Connected libraries

The SharePoint / OneDrive and Google Drive connectors read a library or folder you choose through delegated OAuth granted by a user, with read-only scopes. Tokens are encrypted at rest and revocable at any time. New, updated and deleted files are synced. A connected library can be kept out of automated tagging and embedding so proprietary text is not sent to a model provider unless a user deliberately submits it for testing. Connectors are off by default and enabled on request.

Linking to personas

On a persona, choose Link knowledge base. A persona can read several knowledge bases; a knowledge base can serve many personas. Retiring a knowledge base that personas depend on leaves them ungrounded, so Kumkuat warns before you delete one.

Scheduling

Crawls run on the schedule you set per source. Pausing a knowledge base pauses every source in it.