Purpose
This article introduces the Giotto Python API library, explains how to install and configure it, and shows how to use its main features, including the Responses API, chat-completions compatibility, agents, streaming, and the async client.
Overview
The Giotto Python library provides an OpenAI-style interface for the Giotto API.
It still exposes the older low-level methods, but the recommended entry point is now the higher-level client interface built around:
GiottoClientclient.responses.create(...)client.chat.completions.create(...)client.agents.*
For most users, the preferred way to interact with the API is through the Responses API.
Installation
Install the package
You will be given package wheels (you need Python >= 3.10).
To install the package locally, simply run:
pip install g-client --index-url https://pip-install:gldt-6Hsbe5WSz3ywYXnEMR_s@gitlab.com/api/v4/projects/80522333/packages/pypi/simple
inside your python environment.
Quick start
A minimal example looks like this:
from g_client import GiottoClient
client = GiottoClient(...)
response = client.responses.create(
model="giotto-model",
input="Write a one-sentence bedtime story about a unicorn.",
)
print(response.output_text)
How it works under the hood
When you call:
client.responses.create(...)
the client:
- sends the request to
/orch/orchestrate - receives a
request_id - polls
/orch/fetch/{request_id} - stops when it receives a terminal orchestrator event such as:
request_completedrequest_failed
This means the client provides a simple high-level interface while handling the orchestration polling internally.
Configuration
Required settings
The client needs to know:
- the base API URL
- the Keycloak base URL
- the realm
- the client ID
- the client secret
You can provide these values either:
- explicitly in Python
- through environment variables
If you are using the cloud version, check the API page.
Option 1: Pass credentials explicitly
You can construct the client directly with all configuration values:
from g_client import GiottoClient
client = GiottoClient(
base_url = "<your-base-url>/api",
keycloak_base_url = "<your-base-url>/keycloak",
realm = "giotto",
client_id = "python-client", # or <your-python-client-id> for cloud
client_secret = "<your-client-secret>",
)
Your client secret can be obtained via Keycloak. Please refer to Retrieve the Token for the Giotto Python Client for further details.
On cloud, base_url = "https://cloud.giotto.ai/api" and, keycloak_base_url = "https://cloud.giotto.ai/keycloak".
Option 2: Use environment variables
You can also define the settings through environment variables:
export GIOTTO_BASE_URL="<your-base-url>/api"
export GIOTTO_KEYCLOAK_BASE_URL="<your-base-url>/keycloak"
export GIOTTO_REALM="giotto"
export GIOTTO_CLIENT_ID="python-client"
export GIOTTO_CLIENT_SECRET="<your-client-secret>"
Then create the client normally:
from g_client import GiottoClient
client = GiottoClient()
Security note
Do not commit GIOTTO_CLIENT_SECRET to source control.
If a secret was ever exposed in a real environment, it should be rotated. This article explains how to obtain your GIOTTO_CLIENT_SECRET: Retrieve the Token for the Giotto Python Client.
Responses API
Why use the Responses API
The Responses API is the main high-level text interface of the library.
For most use cases, this is the preferred API surface.
Basic example
from g_client import GiottoClient
client = GiottoClient()
response = client.responses.create(
model="giotto-model",
instructions="You are a concise technical assistant.",
input="How do I check whether a Python object is an instance of a class?",
)
print(response.output_text)
Return value
The returned object is a lightweight response model with OpenAI-style helpers.
You can inspect:
print(response.id)
print(response.status)
print(response.output_text)
print(response.to_dict())
print(response.to_json(indent=2))
Start a request without waiting
If you want to start the request without waiting for completion, use:
response = client.responses.create(
model="giotto-model",
input="hello",
poll=False,
)
print(response.request_id)
Then wait later:
final_response = client.responses.wait(response.request_id)
print(final_response.output_text)
This is useful when you want more control over when you block for the result.
Stream orchestrator events
If you pass stream=True, the client returns an iterator over normalized orchestrator events.
Important: this is not token-by-token streaming. Each yielded item corresponds to one non-empty /orch/fetch/{request_id} response.
This is useful for displaying the agent timeline, for example:
- LLM calls
- handoffs
- tool calls
- tool results
- node returns
- final terminal event
Example:
stream = client.responses.create(
model="giotto-model",
input="Compute 123 * 456 using the Python tool.",
stream=True,
)
for event in stream:
print(event.seq, event.kind, event.service_name, event.notes)
final_response = stream.get_final_response()
print(final_response.output_text)
You can also inspect:
-
stream.eventswhile iterating -
stream.final_responseafter the terminal event
Uploading images and files
responses.create(...) still sends ordinary text-only requests as JSON. When you pass files=..., the client sends the request as multipart/form-data and includes the prompt plus any optional orchestration fields.
response = client.responses.create(
model="giotto-model",
input="Describe the attached image.",
files=["diagram.png"],
priority=2,
generation_params={"temperature": 0.2},
)
print(response.output_text)Each file can be a path, a binary file object, bytes, or a tuple like (filename, content, content_type). generation_params, metadata, and file_descriptors may be passed as Python objects; for multipart requests the client JSON-encodes them into the string fields expected by the API.
Current limitations
The Responses API currently does not aim to support everything.
Currently, the following features are not supported and are out of scope:
- embeddings
- true token-by-token streaming
The supported streaming mode is orchestration-event streaming backed by polling.
Chat Completions Compatibility
Overview
For compatibility with code that still uses chat completions, the library provides a compatible surface.
Example:
from g_client import GiottoClient
client = GiottoClient()
completion = client.chat.completions.create(
model="giotto-model",
messages=[
{"role": "developer", "content": "Talk like a pirate."},
{"role": "user", "content": "How do I check if a Python object is an instance of a class?"},
],
)
print(completion.choices[0].message.content)
How compatibility works
Internally:
- developer/system messages are folded into
instructions - user/assistant messages are serialized into conversation text
- the same orchestrator flow used by
responses.create(...)is then called underneath
This makes it easier to reuse code that still expects a chat-completions interface.
Agents API (NOT FOR CLOUD)
Overview
LLM agents are systems where language models can:
- follow specialized instructions
- invoke tools
- delegate work to other agents
- combine results into a final response
In Giotto, each Agentrepresents a specialized orchestration node with its own instructions, tools, and child handoffs.
For example:
- a Python agent may execute code
- a memory agent may store and retrieve information
- a manager agent may delegate work to specialists
The Agents API allows you to define these orchestration structures directly in Python, generate Giotto YAML configurations, and deploy them to the backend orchestrator.
Define agents in Python (NOT FOR CLOUD)
Example:
from g_client import GiottoClient, Agent, PythonTool, MemoryAddTool, MemorySearchTool
client = GiottoClient()
python_agent = Agent(
name="Python-Specialist",
instructions="""
You are a Python specialist. Use Python only when execution is useful.
Return the result to your parent once you have enough information.
""",
tools=[PythonTool()],
description="Python specialist agent.",
)
memory_agent = Agent(
name="Memory-Agent",
instructions="""
You store and retrieve facts. Use memory-add only for storage tasks and
memory-search only for retrieval tasks.
""",
tools=[MemorySearchTool(), MemoryAddTool()],
description="Memory/cache agent.",
)
manager = Agent(
name="manager-llm",
instructions="""
You are the root manager. Delegate only to direct children. Finalize when
the user request is complete.
""",
handoffs=[python_agent, memory_agent],
description="Root manager/controller that delegates tasks to child agents.",
serve_overrides={
"min_replicas": 1,
"max_replicas": 1,
"num_replicas": 1,
"tensor_parallel_size": 1,
"pipeline_parallel_size": 1,
"max_model_len": 10000,
"max_num_batched_tokens": 10000,
},
)
In the final agent, note the specific serve configurations. In case your GPUs do not have enough memory, you can shard your model, decide the context window size and even the scalability properties (with replicas). Again, it is important to estimate (at least roughly) how the model would fit in your GPUs. For example, a 27b model - like the default giotto-model - will need to be shared over 4 x L4 GPUs and have a limited context maxing to 5000 tokens. On an B200/H200, on the other hand, the model can easily fit with a context of 200k tokens. On an H100 the model fits but the context is limited to 64k tokens.
Generate backend YAML (NOT FOR CLOUD)
You can generate the Giotto backend YAML from the Python agent definitions:
yaml_text = client.agents.to_yaml(
name="orchestrator-demo-big",
root_agent=manager,
max_steps=20,
max_try=3,
)
print(yaml_text)
This generates the Giotto orchestrator YAML configuration, including:
nameroot_nodemax_stepsmax_tryservice_registryray_serve_config
If your handoffs form a tree structure, the generated service registry mirrors the same structure. Each agent receives children equal to its handoff agents plus its tools.
Deploy an agent topology (NOT FOR CLOUD)
You can deploy directly from Python-defined agents:
client.agents.deploy(
name="orchestrator-demo-big",
root_agent=manager,
)
This sends the generated YAML to /orch/deploy with:
Content-Type: application/x-yaml
You can also deploy:
- an existing YAML file
- an existing Python dictionary
Deploy from a YAML file (NOT FOR CLOUD)
client.agents.deploy(
name="orchestrator-demo-big",
file="configs/orchestrator-demo-big.yaml",
)
Deploy from a Python dictionary (NOT FOR CLOUD)
client.agents.deploy(
name="orchestrator-demo-big",
payload=config,
)
Run an agent
After the orchestrator is deployed, you can run the agent using:
result = client.agents.run(
manager,
input="Compute 17 * 23 using the Python tool.",
)
print(result.output_text)
This is a convenience wrapper over client.responses.create(...).
The backend still determines the actual routing using:
- the deployed YAML
- the generated system prompts
Custom MCP tools
The library also supports defining custom tools. Custom tool support is still evolving and may change in future releases.
Example:
from g_client import MCPTool
class SlackTool(MCPTool):
def __init__(
self,
*,
name: str = "tool-slack",
mcp_url: str = "http://172.18.0.1:18745/_/slack/mcp",
) -> None:
super().__init__(
name=name,
mcp_url=mcp_url,
description="MCP tool server with connectivity to Slack to send and read messages.",
)
Then attach it to an agent:
weather_agent = Agent(
name="weather-agent",
instructions="Use the Slack tool when the user asks for it.",
tools=[SlackTool()],
)
Async API
Overview
The library also provides an asynchronous client designed to feel similar to OpenAI’s AsyncOpenAI style.
The async client is:
AsyncGiottoClient
Basic async example
import asyncio
from g_client import AsyncGiottoClient
async def main() -> None:
async with AsyncGiottoClient() as client:
response = await client.responses.create(
model="giotto-model",
input="Write a one-sentence bedtime story about a unicorn.",
)
print(response.output_text)
asyncio.run(main())
Async Responses
You can use the async client with the Responses API directly:
from g_client import AsyncGiottoClient
client = AsyncGiottoClient()
response = await client.responses.create(
model="giotto-model",
instructions="You are a concise technical assistant.",
input="How do I check whether a Python object is an instance of a class?",
)
print(response.output_text)
await client.aclose()
Start async without waiting
You can also start a request and wait later:
response = await client.responses.create(
model="giotto-model",
input="hello",
poll=False,
)
final_response = await client.responses.wait(response.request_id)
print(final_response.output_text)
Async streaming orchestrator events
Async streaming works similarly to the sync version:
stream = await client.responses.create(
model="giotto-model",
input="Compute 123 * 456 using the Python tool.",
stream=True,
)
async for event in stream:
print(event.seq, event.kind, event.service_name, event.notes)
final_response = await stream.get_final_response()
print(final_response.output_text)
Async chat-completions compatibility
Example:
completion = await client.chat.completions.create(
model="giotto-orchestrator",
messages=[
{"role": "developer", "content": "Talk like a pirate."},
{"role": "user", "content": "How do I check if a Python object is an instance of a class?"},
],
)
print(completion.choices[0].message.content)
Async Agents (NOT FOR CLOUD)
Agent definitions remain plain synchronous Python objects, but the network operations are async.
Example:
from g_client import Agent, AsyncGiottoClient, PythonTool
python_agent = Agent(
name="Python-Specialist",
instructions="You are a Python specialist.",
tools=[PythonTool()],
)
manager = Agent(
name="manager-llm",
instructions="Delegate to specialists and finalize when complete.",
handoffs=[python_agent],
)
async with AsyncGiottoClient() as client:
await client.agents.deploy(
name="orchestrator-demo",
root_agent=manager,
)
result = await client.agents.run(
manager,
input="Compute 17 * 23 using the Python specialist.",
)
print(result.output_text)
Summary
The Giotto Python API library gives you a higher-level, OpenAI-style way to interact with Giotto.
Its main strengths are:
- a clean Responses API
- compatibility with chat completions
- support for orchestrator-event streaming
- a code-first Agents API
- both sync and async client surfaces
For most users, the best starting point is:
- install the package
- configure credentials
- try
client.responses.create(...) - move on to streaming, chat compatibility, or agents as needed
Comments
0 comments
Article is closed for comments.