Skip to content

LLM support

The DialoX platform has builtin support for ChatGPT and other large language models (LLMs).

Prompt example

By creating a script of the LLM Prompt, it is possible to define one or more prompts that can be used at runtime in the bot.

A prompt file looks at minimum like this:

prompts:
  - id: rhyme
    text: |
      make a sentence that rhymes with: {{ text }}

In bubblescript this exposes a constant called @prompts.rhyme, that can then be used like this:

dialog main do
  ask "Enter a sentence and I will make it rhyme for you"

  _result = LLM.complete(@prompts.rhyme, text: answer.text)
  say _result.text
end

resulting in a conversation like this:

bot:  Enter a sentence and I will make it rhyme for you
user: I want to fly away!
bot:  Today is Sunday Funday, let's go play!

Individual Prompt Files

As an alternative to managing all prompts in a single prompts.yaml file, the platform supports creating individual prompt files. This approach is particularly useful for bots with many prompts, as it provides better organization and makes it easier to manage prompts individually.

Individual prompt files are stored as markdown with YAML front matter and are configured through the CMS system using the type: prompt content definition. See Individual Prompt Files (type: prompt) in the CMS documentation for detailed configuration instructions.

Both approaches can coexist in the same bot:

  • Prompts defined in prompts.yaml are available as @prompts.id
  • Prompts defined as individual files are also available as @prompts.id
  • All prompts are merged into the same @prompts constant

This means you can use both methods in the same bot, and reference them the same way in your Bubblescript code. Just ensure that prompt IDs are unique across both approaches to avoid conflicts.

Prompt parameters

Only id and text are required. The rest depend on the provider, model, and what you need the call to do.

id

Unique identifier. Exposed in Bubblescript as @prompts.[id].

id: summarize

label

Human-readable name, used when the prompt is included in the CMS or Inbox widget.

label: Summarize

provider

Which LLM backend to call:

  • microsoft_openai — Microsoft Azure OpenAI (default). EU regions.
  • openai — same Azure endpoint; kept as an alias
  • direct_openai — US-based OpenAI (api.openai.com)
  • google_ai — Google Vertex AI / Gemini
provider: microsoft_openai

model

Model id for the chosen provider. Leave it empty to use the platform default (medium).

small, medium, and large resolve per provider:

  • Azure / OpenAI: gpt-4.1-nano, gpt-4.1-mini, gpt-4.1
  • Google: gemini-2.5-flash-lite, gemini-2.5-flash, gemini-2.5-pro

Current models:

Azure OpenAI (microsoft_openai / openai) and US OpenAI (direct_openai):

  • gpt-4.1, gpt-4.1-mini, gpt-4.1-nano
  • gpt-5-mini, gpt-5-nano
  • gpt-4o, gpt-4o-mini

Google (google_ai):

  • gemini-2.5-flash, gemini-2.5-flash-lite, gemini-2.5-pro
  • gemini-3.5-flash, gemini-3.5-flash-lite, gemini-3.6-flash, gemini-3.7-flash, gemini-3.8-flash
  • gemini-3-flash-preview, gemini-3-pro-preview

Older gpt-4, gpt-3.5-turbo, and Gemini 1.5 / 2.0 ids still validate but are deprecated.

model: gpt-4.1-mini

text

Prompt body. A string or a $i18n map. Liquid template: {{ }} bindings are filled by LLM.complete(). Lines starting with system:, user:, or assistant: become separate chat messages.

text:
  $i18n: true
  nl: |
    system: Gegeven de volgende tekst, maak een korte samenvatting.
    user: {{ text }}
  en: |
    system: Given the following text, create a short summary.
    user: {{ text }}

endpoint_params

Extra keys merged onto the provider request body. Use this for parameters that are not first-class prompt fields. The merge is shallow, so a nested key like generationConfig replaces the one the provider built.

endpoint_params:
  some_extra_param: 1

response_format

text, json_object, or json_schema.

response_format: text

response_json_schema

Used only when response_format is json_schema. A subset of JSON Schema; see OpenAI structured outputs.

response_json_schema:
  name: "My schema"
  schema:
    type: object
    properties:
      summary:
        type: string
        description: "A concise summary of the input text"

logprobs

When true, the provider returns log probabilities for generated tokens. They land in _result.raw, not on _result itself. The YAML accepts this field for any model; the provider API rejects it when unsupported.

Works on gpt-4.1, gpt-4.1-mini, gpt-4.1-nano, gpt-4o, and gpt-4o-mini (Azure and US OpenAI). Reasoning models (gpt-5-mini, gpt-5-nano) do not support it.

On Google it is sent as responseLogprobs. That works on Gemini 2.x. Deprecated on Gemini 3.x.

logprobs: false

max_completion_tokens

Maximum tokens in the completion. Sent as max_completion_tokens to OpenAI (not max_tokens).

max_completion_tokens: 100

candidate_count

How many alternative completions to generate (1–10). _result.text still holds only the first; extras, if the provider sent them, are in _result.raw. The YAML accepts this field for any model; the provider API rejects it when unsupported.

OpenAI/Azure send this as n. gpt-4.1, gpt-4.1-mini, gpt-4.1-nano, gpt-4o, and gpt-4o-mini accept it.

Google sends it as candidateCount. Gemini 2.x allows more than one (usually 1–8). Gemini 3.x generally allows only 1.

candidate_count: 1

frequency_penalty

Token frequency penalty, −2.0 to 2.0.

frequency_penalty: 0.0

presence_penalty

Token presence penalty, −2.0 to 2.0.

presence_penalty: 0.0

seed

Optional seed for more deterministic completions.

seed: 42

temperature

Randomness of the output, 0.0 to 2.0. Values outside that range are rejected. For models that do not accept temperature (gpt-5-mini and gpt-5-nano are examples), the field is ignored and a warning is emitted.

temperature: 1.0

top_p

Nucleus sampling parameter.

top_p: 1.0

reasoning_effort

Optional. The YAML schema accepts none, minimal, low, medium, high, xhigh, and max. The allowed subset is validated via llm_db for the selected provider + model.

Gemini 3.x models that expose an effort enum accept this field (mapped to thinkingConfig.thinkingLevel). Gemini 2.5 models use a numeric thinking budget instead, so reasoning_effort is rejected — set thinkingConfig.thinkingBudget through endpoint_params. gemini-3-pro-preview has no effort list in llm_db, so the field is rejected there as well.

gpt-5-mini accepts reasoning_effort. Temperature on that model is ignored:

prompts:
  - id: classify
    model: gpt-5-mini
    reasoning_effort: low
    response_format: json_schema
    response_json_schema:
      name: classification
      schema:
        type: object
        properties:
          intent:
            type: string
            enum: [order, support, other]
          confidence:
            type: number
        required: [intent]
    text: |
      system: Classify the user message. Return only the schema.
      user: {{ text }}
dialog main do
  ask "How can I help?"

  _result = LLM.complete(@prompts.classify, text: answer.text)
  say "I think this is a #{_result.json.intent} request"
end

Raise effort for one call without editing the YAML:

dialog handle_complaint do
  _prompt = LLM.prompt(@prompts.classify, reasoning_effort: "high")
  _result = LLM.complete(_prompt, text: answer.text)
  say _result.json.intent
end

On Gemini 2.5, pass the budget through endpoint_params. Those keys merge at the top of the request body, so a generationConfig here replaces the one the provider built:

prompts:
  - id: research
    provider: google_ai
    model: gemini-2.5-flash
    endpoint_params:
      generationConfig:
        thinkingConfig:
          thinkingBudget: 2048
    text: |
      system: Answer from the conversation so far.
      [[ transcript ]]

stop

Optional list of sequences where the API should stop generating.

stop:
  - "END"

request_timeout

Optional timeout in seconds. On timeout, finish_reason is "timeout".

request_timeout: 30

tools

Optional tools the model may call. Task tools are documented in Tool calling.

tools:
  - type: function
    function:
      name: get_current_weather
      description: Get the current weather in a given location
      parameters:
        type: object
        properties:
          location:
            type: string
          unit:
            type: string
            enum: [celsius, fahrenheit]
        required: [location]

Executing prompts

LLM.complete(prompt, bindings) sends the prompt to the provider. The prompt is usually a constant from a prompt YAML file, such as @prompts.summarize. Bindings fill Liquid variables in the prompt text.

dialog summarize_ticket do
  ask "Paste the ticket"
  _result = LLM.complete(@prompts.summarize, text: answer.text)

  say _result.text
  log "tokens=#{_result.usage} ms=#{_result.request_time}"
end

The full result of the LLM.complete call is a map array which contains the following:

  • text - The output text that LLM produced
  • json - A JSON deserialized version of the text; the runtime detects whether JSON is available in the result and, if so, parses it. The JSON message itself can be padded with arbitrary other texts.
  • usage - The total tokens that were used for this API call
  • request_time - The nr of milliseconds this request took
  • raw - The raw OpenAPI response

User / bot / assistant roles

The prompt text can contain user:, assistant: or system: strings, which will be used for determining the different parts of the prompt (e.g. constructing the messages part of the OpenAPI request payload).

Automatic bindings

Some prompt bindings are done automatically.

In the case of Bubblescript LLM.complete calls, the following bindings are filled automatically:

  • locale - The conversation's locale
  • transcript - The last 5 turns of the bot / user. This is typically used to make a generic chatbot that responds to the previous conversation in a natural way.
  • full_transcript - The full transcript of the conversation. Both transcript and full_transcript are in a format that can be used direclty in the OpenAPI-compatible request payload format, including the role field and tool call messages.

Long conversations can be compacted with LLM.compact() so later LLM.complete calls do not send the entire history.

Compacting full_transcript

LLM.compact() summarizes the conversation so far through the overridable compact_full_transcript platform prompt and stores a compaction annotation. After compaction, full_transcript is rebuilt as:

  1. System messages from before the compaction
  2. The compaction message (a system recap)
  3. Turns that happened after the compaction

Messages before the compaction (except those system messages) are dropped. Bots can override the platform prompt by defining a prompt with the same id, or pass a prompt explicitly: LLM.compact(@prompts.my_compact).

dialog main do
  LLM.compact()
  _result = LLM.complete(@prompts.agent)
  say _result.text
end
  • bot - The metadata of the bot, for instance {{ bot.title }} is exposed.
  • conversation - Some metadata of the conversation, like addr, tags, frontend.
  • user - The contact of this converstaion.
  • persona - The persona as filled in on the "AI" > "Persona" page in a bot. You can also manually construct the persona using the bot binding (bot.purpose, bot.extended_purpose, and bot.guardrails).
  • constants - All constants (for example, @foo "bar") from the bot.
  • prompt_constants - Same as the constants, but these can be rendered as sub-templates for the prompt. For example: {{ prompt_constants.prompt_customizations }}.

The transcript and full_transcript bindings are array bindings and needs to be specified as [[ transcript ]] or [[ full_transcript ]] so with square brackets, and on a line by itself!

prompts:
  - id: agent
    text: |
      system: You are {{ bot.title }}. Speak in locale {{ locale }}.
      [[ transcript ]]
dialog main do
  _result = LLM.complete(@prompts.agent)
  say _result.text
end

No extra bindings are required here: locale, bot, and transcript are filled from the conversation.

Any variables in the Liquid template that were not passed explicitly in the LLM.complete call will be lookup up automatically in the globals of the conversation.

Charging

For every LLM.complete call, a charge event (of type llm.complete) is created and is taken into account in the customer's billing cycle.