Sitecore Collection
A Sitecore Collection enables SearchBlox to connect to your Sitecore CMS instance and index published pages and content for search.
Once configured, users can search Sitecore content directly from the SearchBlox search interface without needing to switch to Sitecore.
Before You Begin
Before configuring a Sitecore Collection, ensure you have the following:
-
The Base URL of your Sitecore instance (for example,
https://your-sitecore.com). -
A Sitecore username and password with read access to the content you want to index.
-
The Root Item ID (GUID) of the Sitecore item where crawling should begin. This can be found in the Sitecore Content Editor under Item Properties.
-
The name of the Sitecore database to crawl:
- web – for published content.
- master – for draft and unpublished content.
-
The language code for the content you want to index (for example,
enfor English oresfor Spanish).
Creating a Sitecore Collection
You can create a Sitecore Collection by following these steps:
-
Log in to the Admin Console, go to the Collections tab, and click Create or the “+” icon.
-
Select Sitecore Collection as the collection type.
-
Enter a unique name for the collection. The name must contain 3–36 alphanumeric characters, and only underscores (_) are allowed.
-
Enable or disable RAG (Retrieval Augmented Generation) depending on your requirement. Enable it if the collection will be used for AI-powered search or chatbot responses.
-
Enable Knowledge Graph if you want SearchBlox to extract entities and relationships from the documents in the collection.
-
Choose whether the collection should be Private or Public. Enable Private Collection Access to restrict the collection to authenticated users only.
-
Configure Collection Encryption if you want to encrypt document content or specific metadata fields.
-
Select the Collection Language based on the language used in the documents. The default language is English.
-
Click Create to create the Sitecore Collection.
-
After the collection is created, you will be redirected to the Sitecore Settings / Authentication section to configure the connection and access details.

Configuring Sitecore Settings
To configure Sitecore integration for your collection, follow these steps:
-
Go to the Sitecore Credentials tab within the collection.
-
Enter the Base URL.
This is the URL of your Sitecore instance (e.g.,https://your-sitecore.com) used to connect and retrieve content. -
Enter the Username.
Provide the Sitecore username required for authentication. -
Enter the Password.
Provide the corresponding password for the Sitecore user to securely access the instance. -
Enter the API Key (Optional).
If your Sitecore setup requires additional authentication, provide the API key here. -
Enter the Database.
Specify the Sitecore database name from which content should be crawled (e.g.,web,master). -
Enter the Root Item ID.
Provide the GUID of the Sitecore item from where crawling should begin. This defines the starting point of content indexing. -
Enter the Language.
Specify the language code for the content to be indexed (e.g.,enfor English,esfor Spanish). -
Click Save to store the configuration and enable the system to crawl and index content from your Sitecore instance.

Collection Settings
Generate Using LLM
- Enable Title to automatically generate concise and relevant titles for the documents using LLM while indexing.
- Enable Description to generate relevant descriptions for the documents using LLM while indexing.
- Enable Topics to generate relevant topics for the documents using LLM while indexing.
Process Images Using LLM
- Enable Generate Description to extract images from documents and generate descriptions using LLM while indexing.
- Enable Enable OCR to extract text from images and scanned documents using the vision model while indexing.
Relevance
- Auto Relevance — Enable to use Hybrid Search for automatic relevance ranking. Compare Keyword Search with Hybrid to evaluate the difference before enabling.
RAG
- Enable RAG — Turn ON to build a retrieval index (vector embeddings) for this collection so it can be used for RAG, Hybrid Search, and AI Answers. This setting applies at the next index — re-index the collection to build or rebuild the RAG index.
RAG Chunking
-
Chunking Strategy — Determines how documents are split for the RAG vector index. Default (recursive) suits most content.
-
Size Unit — Measure chunk size and overlap using one of three modes:
- Global Default — Inherits the system-wide default chunking configuration.
- Tokens — Measure chunk size and overlap in tokens.
- Characters — Measure chunk size and overlap in characters.
-
Chunk Size — Leave blank to inherit the global default (300 tokens).
-
Overlap — Overlap between adjacent chunks; must be less than the chunk size. Leave blank to inherit the global default (30).
Note: Chunking applies at indexing — set it before the first index. Changing it later requires re-indexing the collection.
Click Save to apply the RAG Chunking configuration, or click Cancel/Save at the bottom of the page to discard or store all Collection Settings.
Raw HTML Storage
- Store Raw HTML — Enable to keep the raw fetched/rendered HTML of each crawled page in a per-collection store, for review and preview. Applies to HTML pages only.
Click Save to store the collection settings, or Cancel to discard changes.


Synonyms
Synonyms help the search show relevant documents even when the exact search word is not used.
For example, if someone searches for “global,” the results can also include documents that use “world” or “international.”
We have an option to load Synonyms from the existing documents.

Stopwords
Stopwords are common, high-frequency words that carry minimal semantic value and are typically excluded during text processing, indexing, or search operations. Examples include articles (a, an, the), conjunctions (and, but, or), prepositions (in, on, about), and auxiliary verbs (is, was, would, had).
Purpose:
Reduce noise in search indexing and text analysis
Improve processing efficiency by excluding low-value tokens
Enhance search relevance by prioritizing meaningful keywords

Schedule and Index
Sets the frequency and the start date/time for indexing a collection. Schedule Frequency supported in SearchBlox is as follows:
- Once
- Hourly
- Daily
- Every 48 Hours
- Every 96 Hours
- Weekly
- Monthly
The following operation can be performed in Azure blob collections
| Activity | Description |
|---|---|
| Enable Scheduler for Indexing | Once enabled, you can set the Start Date and Frequency |
| Schedule | For each collection, indexing can be scheduled based on the above options. |
| View all Schedules | Redirects to the Schedules section, where all the Collection Schedules are listed. |
Manage Documents Tab
-
Using Manage Documents tab we can do the following operations:
- Filter
- View content
- View metadata
- Refresh
- Delete
-
To delete a file from your collection, enter the file path and click "Delete".
-
To see the status of an indexed file, click "View Metadata".

Data Fields Tab
Using the Data Fields tab, you can create custom fields for search and view the default and configured fields for the collection.
- Toggle Show Defaults ON to display the collection's default system fields, in addition to any custom fields.
- Use the + icon to add a new custom Data Field.
- Use the info icon to view details about field configuration.
- Use the refresh icon to reload the fields list.
Each field is listed with the following columns:
| Column | Description |
|---|---|
| Name | The name of the data field (e.g., col_id, content_suggest, topics). |
| Type | The data type assigned to the field. |
| Analyzer | The text analyzer applied to the field, if any (e.g., comma_analyzer). Shown as — when no analyzer is applied. |
SearchBlox supports the following Data Field types:
| Type | Description |
|---|---|
| Keyword | Used for alphanumeric values such as IDs, tags, codes, or other exact-match fields (e.g., col_id, faq_content, image_path). |
| Text | Used for full-text search within custom field content (e.g., content_suggest, topics). |
| KNN_Vector | Used to store vector embeddings for semantic/similarity search (e.g., page_dna_vector). |
| Binary | Used to store binary data such as images or files (e.g., imagedata). |
| Boolean | Used for true/false values (e.g., needsReview, approved). |
| Number | Used for numeric values such as prices, quantities, ratings, or counts. |
| Date | Used for date values that can be searched, sorted, and filtered. |
Note: Once Data Fields are configured, the collection must be cleared and re-indexed for the changes to take effect.

Prompts
- When LLM/RAG is enabled, you can edit AI-based prompts for Title, Description, Topic, Image Description, and Smart FAQs.
- You can customize these prompts anytime, and use Restore Default to reset them back to the original SearchBlox settings.


Sitecore Collection Models
The Models page allows you to configure and override AI models used for embeddings, reranking, and LLM-based features within the collection.

Embedding
- Provider specifies the embedding provider used to generate vector representations of documents.
- Model defines the embedding model used to convert document content into vectors for semantic search.
Reranker
- Provider specifies the reranker provider used for improving search result relevance.
- Model defines the reranker model used to re-score and reorder search results based on relevance.
LLM
-
Provider specifies the Large Language Model provider used for AI-powered features.
-
Model defines the LLM used for tasks such as document enrichment, summaries, and SmartFAQs.
-
These settings override global configurations and apply only to the current collection.
Knowledge Graph
- Enable Knowledge Graph — Turn ON to extract entities and relationships from this collection into a Knowledge Graph. This setting applies at the next index — re-index the collection to build or rebuild the graph.

Updated 17 days ago
