LLM support
The DialoX platform has builtin support for ChatGPT and other large language models (LLMs).
Prompt example¶
By creating a script of the LLM Prompt, it is possible to define one or more prompts that can be used at runtime in the bot.
A prompt file looks at minimum like this:
prompts:
- id: rhyme
text: |
make a sentence that rhymes with: {{ text }}
In bubblescript this exposes a constant called @prompts.rhyme, that can then be used like this:
dialog main do
ask "Enter a sentence and I will make it rhyme for you"
_result = LLM.complete(@prompts.rhyme, text: answer.text)
say _result.text
end
resulting in a conversation like this:
bot: Enter a sentence and I will make it rhyme for you
user: I want to fly away!
bot: Today is Sunday Funday, let's go play!
Individual Prompt Files¶
As an alternative to managing all prompts in a single prompts.yaml file, the platform supports creating individual prompt files. This approach is particularly useful for bots with many prompts, as it provides better organization and makes it easier to manage prompts individually.
Individual prompt files are stored as markdown with YAML front matter and are configured through the CMS system using the type: prompt content definition. See Individual Prompt Files (type: prompt) in the CMS documentation for detailed configuration instructions.
Both approaches can coexist in the same bot:
- Prompts defined in
prompts.yamlare available as@prompts.id - Prompts defined as individual files are also available as
@prompts.id - All prompts are merged into the same
@promptsconstant
This means you can use both methods in the same bot, and reference them the same way in your Bubblescript code. Just ensure that prompt IDs are unique across both approaches to avoid conflicts.
Prompt parameters¶
Only id and text are required. The rest depend on the provider, model,
and what you need the call to do.
id¶
Unique identifier. Exposed in Bubblescript as @prompts.[id].
id: summarize
label¶
Human-readable name, used when the prompt is included in the CMS or Inbox widget.
label: Summarize
provider¶
Which LLM backend to call:
microsoft_openai— Microsoft Azure OpenAI (default). EU regions.openai— same Azure endpoint; kept as an aliasdirect_openai— US-based OpenAI (api.openai.com)google_ai— Google Vertex AI / Gemini
provider: microsoft_openai
model¶
Model id for the chosen provider. Leave it empty to use the platform
default (medium).
small, medium, and large resolve per provider:
- Azure / OpenAI:
gpt-4.1-nano,gpt-4.1-mini,gpt-4.1 - Google:
gemini-2.5-flash-lite,gemini-2.5-flash,gemini-2.5-pro
Current models:
Azure OpenAI (microsoft_openai / openai) and US OpenAI (direct_openai):
gpt-4.1,gpt-4.1-mini,gpt-4.1-nanogpt-5-mini,gpt-5-nanogpt-4o,gpt-4o-mini
Google (google_ai):
gemini-2.5-flash,gemini-2.5-flash-lite,gemini-2.5-progemini-3.5-flash,gemini-3.5-flash-lite,gemini-3.6-flash,gemini-3.7-flash,gemini-3.8-flashgemini-3-flash-preview,gemini-3-pro-preview
Older gpt-4, gpt-3.5-turbo, and Gemini 1.5 / 2.0 ids still validate
but are deprecated.
model: gpt-4.1-mini
text¶
Prompt body. A string or a $i18n map. Liquid template: {{ }}
bindings are filled by LLM.complete(). Lines starting with system:,
user:, or assistant: become separate chat messages.
text:
$i18n: true
nl: |
system: Gegeven de volgende tekst, maak een korte samenvatting.
user: {{ text }}
en: |
system: Given the following text, create a short summary.
user: {{ text }}
endpoint_params¶
Extra keys merged onto the provider request body. Use this for
parameters that are not first-class prompt fields. The merge is
shallow, so a nested key like generationConfig replaces the one the
provider built.
endpoint_params:
some_extra_param: 1
response_format¶
text, json_object, or json_schema.
response_format: text
response_json_schema¶
Used only when response_format is json_schema. A subset of JSON
Schema; see OpenAI structured outputs.
response_json_schema:
name: "My schema"
schema:
type: object
properties:
summary:
type: string
description: "A concise summary of the input text"
logprobs¶
When true, the provider returns log probabilities for generated
tokens. They land in _result.raw, not on _result itself. The YAML
accepts this field for any model; the provider API rejects it when
unsupported.
Works on gpt-4.1, gpt-4.1-mini, gpt-4.1-nano, gpt-4o, and
gpt-4o-mini (Azure and US OpenAI). Reasoning models (gpt-5-mini,
gpt-5-nano) do not support it.
On Google it is sent as responseLogprobs. That works on Gemini 2.x.
Deprecated on Gemini 3.x.
logprobs: false
max_completion_tokens¶
Maximum tokens in the completion. Sent as max_completion_tokens to
OpenAI (not max_tokens).
max_completion_tokens: 100
candidate_count¶
How many alternative completions to generate (1–10). _result.text
still holds only the first; extras, if the provider sent them, are in
_result.raw. The YAML accepts this field for any model; the provider
API rejects it when unsupported.
OpenAI/Azure send this as n. gpt-4.1, gpt-4.1-mini, gpt-4.1-nano,
gpt-4o, and gpt-4o-mini accept it.
Google sends it as candidateCount. Gemini 2.x allows more than one
(usually 1–8). Gemini 3.x generally allows only 1.
candidate_count: 1
frequency_penalty¶
Token frequency penalty, −2.0 to 2.0.
frequency_penalty: 0.0
presence_penalty¶
Token presence penalty, −2.0 to 2.0.
presence_penalty: 0.0
seed¶
Optional seed for more deterministic completions.
seed: 42
temperature¶
Randomness of the output, 0.0 to 2.0. Values outside that range are
rejected. For models that do not accept temperature (gpt-5-mini and
gpt-5-nano are examples), the field is ignored and a warning is
emitted.
temperature: 1.0
top_p¶
Nucleus sampling parameter.
top_p: 1.0
reasoning_effort¶
Optional. The YAML schema accepts none, minimal, low, medium,
high, xhigh, and max. The allowed subset is validated via llm_db
for the selected provider + model.
Gemini 3.x models that expose an effort enum accept this field (mapped
to thinkingConfig.thinkingLevel). Gemini 2.5 models use a numeric
thinking budget instead, so reasoning_effort is rejected — set
thinkingConfig.thinkingBudget through endpoint_params.
gemini-3-pro-preview has no effort list in llm_db, so the field is
rejected there as well.
gpt-5-mini accepts reasoning_effort. Temperature on that model is
ignored:
prompts:
- id: classify
model: gpt-5-mini
reasoning_effort: low
response_format: json_schema
response_json_schema:
name: classification
schema:
type: object
properties:
intent:
type: string
enum: [order, support, other]
confidence:
type: number
required: [intent]
text: |
system: Classify the user message. Return only the schema.
user: {{ text }}
dialog main do
ask "How can I help?"
_result = LLM.complete(@prompts.classify, text: answer.text)
say "I think this is a #{_result.json.intent} request"
end
Raise effort for one call without editing the YAML:
dialog handle_complaint do
_prompt = LLM.prompt(@prompts.classify, reasoning_effort: "high")
_result = LLM.complete(_prompt, text: answer.text)
say _result.json.intent
end
On Gemini 2.5, pass the budget through endpoint_params. Those keys
merge at the top of the request body, so a generationConfig here
replaces the one the provider built:
prompts:
- id: research
provider: google_ai
model: gemini-2.5-flash
endpoint_params:
generationConfig:
thinkingConfig:
thinkingBudget: 2048
text: |
system: Answer from the conversation so far.
[[ transcript ]]
stop¶
Optional list of sequences where the API should stop generating.
stop:
- "END"
request_timeout¶
Optional timeout in seconds. On timeout, finish_reason is "timeout".
request_timeout: 30
tools¶
Optional tools the model may call. Task tools are documented in Tool calling.
tools:
- type: function
function:
name: get_current_weather
description: Get the current weather in a given location
parameters:
type: object
properties:
location:
type: string
unit:
type: string
enum: [celsius, fahrenheit]
required: [location]
Executing prompts¶
LLM.complete(prompt, bindings) sends the prompt to the provider. The
prompt is usually a constant from a prompt YAML file, such as
@prompts.summarize. Bindings fill Liquid variables in the prompt text.
dialog summarize_ticket do
ask "Paste the ticket"
_result = LLM.complete(@prompts.summarize, text: answer.text)
say _result.text
log "tokens=#{_result.usage} ms=#{_result.request_time}"
end
The full result of the LLM.complete call is a map array which contains the following:
text- The output text that LLM producedjson- A JSON deserialized version of the text; the runtime detects whether JSON is available in the result and, if so, parses it. The JSON message itself can be padded with arbitrary other texts.usage- The total tokens that were used for this API callrequest_time- The nr of milliseconds this request tookraw- The raw OpenAPI response
User / bot / assistant roles¶
The prompt text can contain user:, assistant: or system: strings,
which will be used for determining the different parts of the prompt
(e.g. constructing the messages part of the OpenAPI request payload).
Automatic bindings¶
Some prompt bindings are done automatically.
In the case of Bubblescript LLM.complete calls, the following bindings
are filled automatically:
locale- The conversation's localetranscript- The last 5 turns of the bot / user. This is typically used to make a generic chatbot that responds to the previous conversation in a natural way.full_transcript- The full transcript of the conversation. Bothtranscriptandfull_transcriptare in a format that can be used direclty in the OpenAPI-compatible request payload format, including therolefield and tool call messages.
Long conversations can be compacted with LLM.compact() so later LLM.complete calls do not send the entire history.
Compacting full_transcript¶
LLM.compact() summarizes the conversation so far through the overridable compact_full_transcript platform prompt and stores a compaction annotation. After compaction, full_transcript is rebuilt as:
- System messages from before the compaction
- The compaction message (a
systemrecap) - Turns that happened after the compaction
Messages before the compaction (except those system messages) are dropped. Bots can override the platform prompt by defining a prompt with the same id, or pass a prompt explicitly: LLM.compact(@prompts.my_compact).
dialog main do
LLM.compact()
_result = LLM.complete(@prompts.agent)
say _result.text
end
bot- The metadata of the bot, for instance{{ bot.title }}is exposed.conversation- Some metadata of the conversation, likeaddr,tags,frontend.user- The contact of this converstaion.persona- The persona as filled in on the "AI" > "Persona" page in a bot. You can also manually construct the persona using thebotbinding (bot.purpose,bot.extended_purpose, andbot.guardrails).constants- All constants (for example,@foo "bar") from the bot.prompt_constants- Same as theconstants, but these can be rendered as sub-templates for the prompt. For example:{{ prompt_constants.prompt_customizations }}.
The
transcriptandfull_transcriptbindings are array bindings and needs to be specified as[[ transcript ]]or[[ full_transcript ]]so with square brackets, and on a line by itself!
prompts:
- id: agent
text: |
system: You are {{ bot.title }}. Speak in locale {{ locale }}.
[[ transcript ]]
dialog main do
_result = LLM.complete(@prompts.agent)
say _result.text
end
No extra bindings are required here: locale, bot, and transcript
are filled from the conversation.
Any variables in the Liquid template that were not passed explicitly in the LLM.complete call will be lookup up automatically in the globals of the conversation.
Charging¶
For every LLM.complete call, a charge event (of type llm.complete) is
created and is taken into account in the customer's billing cycle.