> ## Documentation Index
> Fetch the complete documentation index at: https://docs.qontext.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Website

> Crawl and ingest content from any public website into your context repository

## Overview

The Website connection allows Qontext to crawl public web pages and ingest their text content into your context repository.

**Supported content:**

| Crawl type    | Description                                                                           |
| ------------- | ------------------------------------------------------------------------------------- |
| `Single Page` | Crawls only the specific URL provided. Best for individual articles or pages          |
| `Deep Crawl`  | Crawls the URL and follows all internal links. Best for entire sites or documentation |

**Unsupported:**

| Content            | Description                                                                   |
| ------------------ | ----------------------------------------------------------------------------- |
| `Videos`           | Video files and embedded video content                                        |
| `Restricted sites` | Some websites (e.g. LinkedIn) block automated crawling and cannot be ingested |

***

## Setting up the connection

You can add a website from your Qontext workspace, enter the URL, choose a crawl type, and set a sync frequency. The steps below walk you through the full flow.

<Steps>
  <Step title="Open Sources in your workspace">
    In the [Qontext app](https://app.qontext.ai), open your workspace and go to the **Sources** tab, then click **+ Connect data source**.

    <img src="https://mintcdn.com/qontext/j3o0tEnncO_MaI30/connections/images/sources_page.png?fit=max&auto=format&n=j3o0tEnncO_MaI30&q=85&s=11648406c2e1aef5cf37535482067443" alt="Sources tab in a Qontext workspace" className="rounded-lg mt-4" style={{ maxWidth: '100%' }} width="2977" height="1856" data-path="connections/images/sources_page.png" />
  </Step>

  <Step title="Select Website as the data source">
    Choose **Website** from the list of available data sources. This starts the website connection flow.

    <img src="https://mintcdn.com/qontext/QPEmfEYHnv00ykRo/connections/images/qontext_select_integration.png?fit=max&auto=format&n=QPEmfEYHnv00ykRo&q=85&s=333cc38e25e2a6c6003ebe5a0b64d492" alt="Select Website from the list of integrations" className="rounded-lg mt-4" style={{ maxWidth: '100%' }} width="1502" height="1075" data-path="connections/images/qontext_select_integration.png" />
  </Step>

  <Step title="Enter the website URL">
    Give the website a **name**, which will be displayed on the Sources table. Then, enter the **website URL** you want to crawl. HTTPS is assumed by default; for HTTP sites, enter the full URL.

    URLs used in other workspaces are shown for reference. URLs already connected to this workspace are displayed and cannot be connected again, but the crawl type of an already ingested website can be edited in the **Sources** tab.

    <img src="https://mintcdn.com/qontext/QPEmfEYHnv00ykRo/connections/images/websites/websites_select_page.png?fit=max&auto=format&n=QPEmfEYHnv00ykRo&q=85&s=5dd0bb7ed99c3c9419f11eaba0afcb10" alt="Add a website URL and name" className="rounded-lg mt-4" style={{ maxWidth: '100%' }} width="1503" height="1094" data-path="connections/images/websites/websites_select_page.png" />
  </Step>

  <Step title="Choose crawl type" id="configure-what-to-sync-filters">
    Select how the website should be crawled. **Single Page** crawls only the specific URL provided and is best for individual articles or pages. **Deep Crawl** crawls the URL and follows all internal links, making it best for entire sites or documentation.

    <img src="https://mintcdn.com/qontext/QPEmfEYHnv00ykRo/connections/images/websites/websites_select_crawl_type.png?fit=max&auto=format&n=QPEmfEYHnv00ykRo&q=85&s=cda62fcfbbffdca0b53535347bd4d6ac" alt="Choose between Single Page and Deep Crawl" className="rounded-lg mt-4" style={{ maxWidth: '100%' }} width="1513" height="768" data-path="connections/images/websites/websites_select_crawl_type.png" />
  </Step>

  <Step title="Set sync frequency">
    Choose how often Qontext should re-crawl the website (e.g. daily, weekly, monthly). More frequent syncs keep the context repository up to date but use more resources.

    <img src="https://mintcdn.com/qontext/QPEmfEYHnv00ykRo/connections/images/websites/websites_sync_frequency.png?fit=max&auto=format&n=QPEmfEYHnv00ykRo&q=85&s=93ff32f099af3ae3ec91fa85fe274186" alt="Sync frequency settings" className="rounded-lg mt-4" style={{ maxWidth: '100%' }} width="754" height="384" data-path="connections/images/websites/websites_sync_frequency.png" />
  </Step>

  <Step title="Review and create connection">
    Review your choices (name, URL, crawl type, frequency), then click **Complete** to create the connection. Qontext will start the initial crawl shortly after.

    <img src="https://mintcdn.com/qontext/QPEmfEYHnv00ykRo/connections/images/websites/websites_review_and_create.png?fit=max&auto=format&n=QPEmfEYHnv00ykRo&q=85&s=bc7d9b3c41c83e37e11a723fe1cdcfa5" alt="Review summary and create website connection" className="rounded-lg mt-4" style={{ maxWidth: '100%' }} width="1497" height="849" data-path="connections/images/websites/websites_review_and_create.png" />
  </Step>
</Steps>

***

## Sync latency

The crawl time depends on the crawl type and the size of the website. Most can be ingested within minutes.

***

## FAQ

<AccordionGroup>
  <Accordion title="When should I use Single Page vs Deep Crawl?">
    Use **Single Page** when you want to ingest a specific article, blog post, or landing page. Use **Deep Crawl** when you want to ingest an entire site or documentation portal. Qontext will follow all internal links starting from the URL you provide.
  </Accordion>

  <Accordion title="How many pages does a Deep Crawl include?">
    A Deep Crawl follows up to 100 pages starting from the provided URL. If you need to limit the scope, consider using Single Page for specific URLs instead.
  </Accordion>

  <Accordion title="Can I add multiple websites to the same workspace?">
    Yes. You can add as many website connections as you need from the **Sources** tab. Each website is a separate data source with its own crawl type and sync schedule.
  </Accordion>

  <Accordion title="How do I manually refresh website data?">
    Manual refresh for website connections is available via the **Sources** tab. Navigate to the respective website connection and open the **Syncs** tab. In the Sync settings, you can select **Resync all data**.

    <img src="https://mintcdn.com/qontext/j3o0tEnncO_MaI30/connections/images/websites/websites_resync.png?fit=max&auto=format&n=j3o0tEnncO_MaI30&q=85&s=437ffbd1799a7287494b86afacbc5d3f" alt="Resync" className="rounded-lg mt-4" style={{ maxWidth: '100%' }} width="2996" height="1860" data-path="connections/images/websites/websites_resync.png" />
  </Accordion>

  <Accordion title="What happens if a page is removed from the website?">
    If a previously crawled page is no longer accessible, it will not be updated during the next sync. Content already ingested remains in your context repository. For data removal, contact [support@qontext.ai](mailto:support@qontext.ai).
  </Accordion>

  <Accordion title="Is the crawled content kept up to date?">
    Qontext re-crawls the website based on the sync frequency you set (daily, weekly, or monthly). Each sync picks up the newest version of the pages and discovers new pages (for Deep Crawl).
  </Accordion>
</AccordionGroup>
