OpenAI-Compatible Endpoint

Direct access to the underlying LLM via an OpenAI-compatible endpoint (/v1/chat/completions, streaming supported), backed by SearchAI Native Inference (or whichever provider is console-active) for client-side and MCP tool-calling integrations.

Available in: SearchBlox V12.2.2

URL

https://localhost:8443/rest/v2/api/ai/openai/v1/chat/completions

Note: localhost:8443 is just the local reference config. Replace it with your own console/server domain (e.g. https://your-searchblox-server.com).

Method
POST

Media Type
application/json

Headers

HeaderRequiredRequired
Content-TypeYesapplication/json
SB-PKEYYesYour SearchBlox user private key

Request Fields

FieldDescriptionTypeRequired
messagesConversation history (system / user / assistant / tool roles).arrayYes
modelModel id to use. Omit to use the console-active model. Available on the SearchBlox server: q35-4b, gemma-4-E4B-it, qwen38-27b.stringNo
streamReturn Server-Sent Events instead of a single JSON body.booleanNo
toolsOpenAI-style function tools. Client-executed - the endpoint returns tool_calls and your client runs them, replying as role:"tool".arrayNo
searchblox_toolsNames of built-in SearchBlox tools (search, collections, KG). Server-executed - the endpoint runs the tool call itself and returns only the final answer.string[]No
mcp_serversMCP servers to expose to the model. Also server-executed.arrayNo

Response Codes

CodeDescription
200Success - completion returned (or SSE stream started).
401Missing or invalid SB-PKEY / Authorization.
400Malformed request body.
413 / context_length_exceededPrompt + tool definitions exceed the active model's max context

Example 1 - Gibberish Classifier

Request:

curl -sk -X POST "https://localhost:8443/rest/v2/api/ai/openai/v1/chat/completions" \
  -H "Content-Type: application/json" -H "SB-PKEY: YOUR_PKEY" \
  -d '{
   "model": "q35-4b",
    "messages": [
      { "role": "system", "content": "You are a strict text classifier. Classify the user's input into exactly one category:\n\n- \"full\": the entire input is meaningless (random keystrokes, no real words, no discernible intent, e.g. \"asdkj wqoeiu\").\n- \"partial\": the input mixes real, meaningful words/phrases with nonsense tokens or gibberish fragments (e.g. \"hello ashjkdfgh\").\n- \"not_gibberish\": the input is coherent, meaningful text with no gibberish.\n\nRules:\n- Judge based on whether recognizable words/intent exist, not spelling errors or typos alone. Minor typos in otherwise real words should NOT count as gibberish.\n- Common internet slang, abbreviations (lol, thx, pls, u, ur) and filler words (blah, meh) count as real, meaningful tokens — not gibberish.\n- A single real word or greeting amid nonsense still counts as \"partial\", not \"not_gibberish\".\n- If the input has no spaces but forms a valid sentence when split at reasonable word boundaries, classify it as \"not_gibberish\".\n- Pure numbers, symbols, or empty input should be classified as \"full\".\n- If unsure between \"partial\" and \"full\", check if ANY substring is a real word in any language — if yes, choose \"partial\".\n- Do not explain your reasoning. Do not add extra fields.\n\nRespond with ONLY a JSON object in this exact form: {\"gibberish\": \"full\" | \"partial\" | \"not_gibberish\"}" },
      { "role": "user", "content": "dsfgds dsdfgdd" }
    ]
  }'

Response:

{
  "choices": [
        {
            "index": 0,
            "message": {
                "role": "assistant",
                "content": "{\"gibberish\": \"full\"}"
            },
            "finish_reason": "stop"
        }
    ]
}

Note: Response field: gibberish

  • "full" - Input is entirely meaningless random keystrokes, symbols, numbers only, or empty input, with no discernible real words or intent.
  • "partial" - Input contains a mix of meaningful words/phrases and nonsense tokens or gibberish fragments.
  • "not_gibberish" - Input is coherent, meaningful text with no gibberish detected.

Example 2 - Tool calling

Mode A (client-executed)

Supply OpenAI tools; the endpoint returns tool_calls. Your client runs the tool and replays results as role:"tool" messages.

Request:

curl -sk -X POST 
"https://localhost:8443/rest/v2/api/ai/openai/v1/chat/completions" \
  -H "Content-Type: application/json" -H "SB-PKEY: YOUR_PKEY" \
  -d '{"messages":[{"role":"user","content":"Weather in Paris?"}],
       "tools":[{"type":"function","function":{"name":"get_weather",
         "parameters":{"type":"object","properties":{"city":{"type":"string"}}}}}]}'
# → response.choices[0].message.tool_calls[...]

Response - the model asks to call the tool instead of answering directly:

{
  "choices": [
    {
      "message": {
        "role": "assistant",
        "tool_calls": [
          {
            "id": "call_1",
            "type": "function",
            "function": { "name": "get_weather", "arguments": "{\"city\":\"Paris\"}" }
          }
        ]
      }
    }
  ]
}

Your client runs get_weather("Paris"), then sends the result back as a role:"tool" message on the same conversation to get the final answer:

curl -sk -X POST 
"https://localhost:8443/rest/v2/api/ai/openai/v1/chat/completions" \
  -H "Content-Type: application/json" -H "SB-PKEY: YOUR_PKEY" \
  -d '{"messages":[
        {"role":"user","content":"Weather in Paris?"},
        {"role":"assistant","tool_calls":[
          {"id":"call_1","type":"function",
           "function":{"name":"get_weather","arguments":"{\"city\":\"Paris\"}"}}
        ]},
        {"role":"tool","tool_call_id":"call_1",
         "content":"{\"temperature\":\"15°C\",\"condition\":\"Cloudy\"}"}
      ],
      "tools":[{"type":"function","function":{"name":"get_weather",
        "parameters":{"type":"object","properties":{"city":{"type":"string"}}}}}]}'

Final response - the model now answers using the tool result you supplied:

{
  "choices": [
    {
      "message": {
        "role": "assistant",
        "content": "It's currently 15°C and cloudy in Paris."
      }
    }
  ]
}

Mode B (server-executed / agentic)

Pass searchblox_tools and/or mcp_servers; the endpoint runs the full agentic loop server-side (via McpToolBridge) and returns only the final answer. This is the way to give a model SearchBlox search/collection/KG tools without wiring an agent.

Request:

curl -sk -X POST 
"https://<your-domain>:8443/rest/v2/api/ai/openai/v1/chat/completions" \
  -H "Content-Type: application/json" -H "SB-PKEY: YOUR_PKEY" \
  -d '{"messages":[{"role":"user","content":"Find information about hybrid search in the searchblox collection."}],
       "searchblox_tools":["searchblox_search"]}'

Response - the tool already ran server-side, so you get the final answer directly, with no tool_calls to handle:

{
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "Based on the searchblox collection, hybrid search combines keyword and semantic retrieval to improve result relevance."
      },
      "finish_reason": "stop"
    }
  ]
}

Did this page help you?