AI Gateway Module
Call Base44’s managed AI models from your own backend code viabase44.asServiceRole.aiGateway.
Note: Backend functions only. It uses your app’s models, billing, and credit quota — there is no API key to manage.
Overview
connection() returns a baseURL, bearer token, and headers to hand to a client
library. Two providers sit behind it:
Every call is metered against your app’s credit quota, the same quota
integrations.Core.* uses.
When to use it
Apps that restrict Core integrations (the default for new apps) also block frontend calls to
GenerateImage, GenerateVideo, and GenerateSpeech; call them from a backend function as
base44.asServiceRole.integrations.Core.*.
Methods
Call it as
base44.asServiceRole.aiGateway.connection(). base44.aiGateway (the
caller’s token) also exists, but new apps restrict Core integrations by default, and a
public app then rejects user-token gateway calls with 403. Service-role calls from a
backend function always pass.
headers requires @base44/sdk 0.8.52+ and provider requires 0.8.50+.
Rules for every gateway call
- Backend function only (
createClientFromRequest(req)). All other backend-function rules (deployment, secrets, error handling) apply — see functions.md. Never send the gatewaytokento the browser. - Guard user-triggered functions with
await base44.auth.me()before the service-role call — every call spends the app owner’s credits. Scheduled or entity-triggered runs have no caller to check. - Decide cost-driving parameters in code (model, image count, video duration, resolution). Don’t pass a browser-supplied request straight to the gateway, or any signed-in user can pick the most expensive options.
- Always pass
headersto the client (defaultHeadersfor theopenaiSDK,headersfor Vercel AI SDK providers). On a client fromcreateClientFromRequest()it carries the signedBase44-Statethat a workspace IP allowlist requires; without it those workspaces reject the call. - Set
maxRetries: 0on image, video, speech, and evaluation calls. Client retries replay billed requests.
Build a code agent
- Get the connection with
base44.asServiceRole.aiGateway.connection()→{ baseURL, token, headers }. - Point an agent SDK’s OpenAI-compatible provider at it (
baseURL,apiKey: token,headers). - Give the agent tools that read/act on your app via
base44.*, and let it finish by recording its result through a tool.
- Tools run in the caller’s scope by default. The gateway connection is service-role,
but the agent’s tools should use
base44.entities.*so they act with the calling user’s permissions (RLS applies) and can’t exceed them. Usebase44.asServiceRole.entities.*in a tool only for genuine cross-user/system work — and then scope it to trusted context, not agent-chosen inputs (e.g. fixcustomer_emailfrom the request, not an agent parameter), since service role has full access. - Stateless between invocations. Persistent memory means storing and replaying state (e.g. in an entity).
- Use model
automaticunless the task needs a specific model — non-default models use more credits: only when needed, and tell the user. - Don’t chain
InvokeLLMto fake a tool loop — use a real agent loop. - Always bound the loop.
stopWhenis an OR-list — the first condition to fire wins (mix a step cap likestepCountIs, a finish tool likehasToolCall, or a custom check). Give it room to finish but stop a runaway: every step is another metered model call.
messages:
baseURL, token, and headers. Streaming (stream: true, streamText) is
supported.
Generate and edit images
POST /images/generations and POST /images/edits through the openai SDK:
- Editing:
client.images.edit({ model, prompt, image })needs at least one input image — uploaded bytes inimage, orreference_image_urls. Generations acceptreference_image_urlstoo. - Other options:
quality,output_format,background(transparent/opaque) — support is model-specific. Withautomatic, setting any of them (evenquality: "auto") skips the Gemini models and routes to a GPT Image model, which changes the output and the cost — omit them unless you need them. - Models: prefer
"automatic", which picks a model that supports the requested options. Pin one (e.g.gemini_3_1_flash_image,gpt_image_2) only when needed. Unsupported option/model combinations return 400. Full model list and limits: Generate images with the AI Gateway. - Cost preview: add
dry_run: trueto the same request to get{ dry_run: true, model, usage: { base44_credits } }without generating or charging.
models.imageModel("automatic") with generateImage and pass
the Base44 extensions under providerOptions: { base44: { ... } } (the key is the name
you gave createOpenAICompatible).
Generate videos
Video generation is an asynchronous job through theopenai SDK’s videos client.
videos.create() returns a job id right away; the video is ready later.
- Statuses:
queued,in_progress,completed,failed. Treatqueued/in_progressor hitting your polling limit as unfinished, not failed. For background flows, persistvideo.id(e.g. in an entity) and retrieve it in a later run. - Cost differs a lot by model. A 4-second clip without audio ranged from 48 credits
(
veo_3_1_lite) to over 350 (kling_3) when this was written. Preview withdry_runbefore choosing a model, and default toveo_3_1_liteunless the app needs another. - Models: there is no
automatic— pass a specific model:veo_3_1_lite,veo_3_1_fast,seedance_2,seedance_2_5,seedance_2_fast,seedance_2_mini,kling_3,minimax_h3,minimax_h3_max,grok_imagine_video,grok_imagine_video_1_5. Supportedseconds,resolution,aspect_ratio, and references differ per model, and some models rejectgenerate_audio; unsupported values return 400, and an unknown model (includingautomatic) returns 404. - Request fields:
model,prompt,seconds,aspect_ratio,resolution,generate_audio,seed,dry_run, and eitherframe_images(up to 2, each{ type: "image_url", image_url: { url }, frame_type: "first_frame" | "last_frame" }) orinput_references(up to 8 image/video/audio URLs, each{ type: "image_url", image_url: { url } }and likewise forvideo_url/audio_url) — don’t combine the two. Reference URLs must be public HTTPS URLs. - Cost preview:
videos.create({ ...request, dry_run: true })returns HTTP 200 withusage.base44_creditsand no job id — don’t poll it, and don’t return it as a 202.
Generate speech
- Browser-only read-aloud, without stored audio or voice/style requirements: use the browser’s TTS API. This option does not apply to native apps.
- Straightforward stored MP3 with a Core voice (
riverdefault,honey,sunny,storm,spark), optional ISO-639-1language_code, up to 5,000 characters: useintegrations.Core.GenerateSpeech({ text, voice?, language_code? })→{ url }. - Model choice, delivery instructions, speed, other formats, or voices outside Core: use the gateway. Follow this section when editing existing gateway speech code too.
POST /audio/speech returns completed audio bytes. It does not store a file or return
a URL or job ID; there is no polling step.
Call it from a backend function, following the
shared rules and functions.md.
text is the supplied spoken text; base44 is the function’s request client:
base44.asServiceRole.integrations.Core.UploadFile({ file }) returns file_url
(see UploadFile).
For private file_uri, call base44.integrations.Core.CreateFileSignedUrl({ file_uri })
and use signed_url for playback (see the signed-URL flow).
Match the filename and content type when choosing another encoding. Prefer MP3 or WAV
for playback; raw PCM requires decoding and cannot be played directly as an audio URL.
- Web: render
<audio controls src={signed_url}>; do not assume autoplay is allowed. - Native: use an installed React Native-compatible player or
Linking.openURL(signed_url)to open saved audio externally. The native template has no audio-player package. Do not use HTML<audio>,new Audio, orspeechSynthesis.
dry_run: true
returns JSON, without generating or charging; the estimate can differ from the final
charge and contains no audio or file to play:
Models and controls: omitted
model or "automatic" resolves to gpt_4o_mini_tts.
- OpenAI legacy voices:
alloy,ash,coral,echo,fable,onyx,nova,sage,shimmer. OpenAI extended addsballad,verse,marin,cedar. - ElevenLabs voices:
honey,river,spark,storm,sunny. - Gemini voices:
Kore,Puck,Zephyr,Charon,Leda,Fenrir(case-sensitive). - OpenAI input limit: 4,096 characters for all three models.
eleven_flash_v2_5language codes:ar,bg,cs,da,de,el,en,es,fi,fil,fr,hi,hr,id,it,ja,ko,ms,nl,pl,pt,ro,ru,sk,sv,ta,tr,uk,zh.
GET /audio/speech/models returns { items, next_cursor };
GET /audio/speech/models/{model} returns one entry with its capabilities. Use these
gateway catalogs for voices, formats and their defaults, speed ranges, instruction and
language support, and input-length limits. The automatic entry’s resolves_to identifies
the fixed default. Do not assume upstream provider features are exposed by the gateway.
AI decisions (Jev)
Structured evaluations: give thejev model some state and a set of questions, and get
each answer back with probabilities. Use it to classify, score, or route; then apply your
own thresholds and actions in deterministic code.
- Needs
ai7.0.105+ — older versions (including the7.0.16code-agent pin above) don’t exportexperimental_evaluate. - The model id is exactly
"jev". The package README’s"jev-latest"is rejected by the gateway.
state— string, JSON object, or array (an array is one state).questions— non-empty; questions can’t depend on other answers in the same call.choice:criteriarequired, an object of 1–255 named options (descriptions may benull).score:criteriarequired, an array of 2–10 ordered rubric levels.boolean:criteriaoptional, an object with onlytrue/falsekeys (descriptions may benull).instructionsand criteria descriptions may be a string, JSON object, or JSON array.
Models
- Chat: any model available through
InvokeLLM;automaticby default (see the code-agent rules above). - Images:
automaticor a pinned image model — see Generate and edit images. - Videos: always a specific model — see Generate videos.
- Speech:
automaticis a fixed default; use the speech catalog for model-specific controls — see Generate speech. - Evaluations:
jev.
Notes
- Token: the service-role token for
base44.asServiceRole.aiGateway; the caller’s token (or an empty string when unauthenticated) forbase44.aiGateway. - Billing: metered per call against your app’s credit quota (same as InvokeLLM). If the app is out of credits, calls are rejected before the model runs.