Skip to main content
This page is part of an AI coding agent skill and is written for agents, not humans. For the human-readable Base44 docs, see the developer documentation.

AI Gateway Module

Call Base44’s managed AI models from your own backend code via base44.asServiceRole.aiGateway.
Note: Backend functions only. It uses your app’s models, billing, and credit quota — there is no API key to manage.

Overview

connection() returns a baseURL, bearer token, and headers to hand to a client library. Two providers sit behind it: Every call is metered against your app’s credit quota, the same quota integrations.Core.* uses.

When to use it

Apps that restrict Core integrations (the default for new apps) also block frontend calls to GenerateImage, GenerateVideo, and GenerateSpeech; call them from a backend function as base44.asServiceRole.integrations.Core.*.

Methods

Call it as base44.asServiceRole.aiGateway.connection(). base44.aiGateway (the caller’s token) also exists, but new apps restrict Core integrations by default, and a public app then rejects user-token gateway calls with 403. Service-role calls from a backend function always pass. headers requires @base44/sdk 0.8.52+ and provider requires 0.8.50+.

Rules for every gateway call

  • Backend function only (createClientFromRequest(req)). All other backend-function rules (deployment, secrets, error handling) apply — see functions.md. Never send the gateway token to the browser.
  • Guard user-triggered functions with await base44.auth.me() before the service-role call — every call spends the app owner’s credits. Scheduled or entity-triggered runs have no caller to check.
  • Decide cost-driving parameters in code (model, image count, video duration, resolution). Don’t pass a browser-supplied request straight to the gateway, or any signed-in user can pick the most expensive options.
  • Always pass headers to the client (defaultHeaders for the openai SDK, headers for Vercel AI SDK providers). On a client from createClientFromRequest() it carries the signed Base44-State that a workspace IP allowlist requires; without it those workspaces reject the call.
  • Set maxRetries: 0 on image, video, speech, and evaluation calls. Client retries replay billed requests.

Build a code agent

  1. Get the connection with base44.asServiceRole.aiGateway.connection() → { baseURL, token, headers }.
  2. Point an agent SDK’s OpenAI-compatible provider at it (baseURL, apiKey: token, headers).
  3. Give the agent tools that read/act on your app via base44.*, and let it finish by recording its result through a tool.
Rules:
  • Tools run in the caller’s scope by default. The gateway connection is service-role, but the agent’s tools should use base44.entities.* so they act with the calling user’s permissions (RLS applies) and can’t exceed them. Use base44.asServiceRole.entities.* in a tool only for genuine cross-user/system work — and then scope it to trusted context, not agent-chosen inputs (e.g. fix customer_email from the request, not an agent parameter), since service role has full access.
  • Stateless between invocations. Persistent memory means storing and replaying state (e.g. in an entity).
  • Use model automatic unless the task needs a specific model — non-default models use more credits: only when needed, and tell the user.
  • Don’t chain InvokeLLM to fake a tool loop — use a real agent loop.
  • Always bound the loop. stopWhen is an OR-list — the first condition to fire wins (mix a step cap like stepCountIs, a finish tool like hasToolCall, or a custom check). Give it room to finish but stop a runaway: every step is another metered model call.
Example with the Vercel AI SDK — a background reviewer the app runs on a return request:
Images as input: when the agent needs to see an image, pass it as an image part in messages:
Any OpenAI-compatible agent SDK works the same way — construct its provider/client with the gateway’s baseURL, token, and headers. Streaming (stream: true, streamText) is supported.

Generate and edit images

POST /images/generations and POST /images/edits through the openai SDK:
  • Editing: client.images.edit({ model, prompt, image }) needs at least one input image — uploaded bytes in image, or reference_image_urls. Generations accept reference_image_urls too.
  • Other options: quality, output_format, background (transparent/opaque) — support is model-specific. With automatic, setting any of them (even quality: "auto") skips the Gemini models and routes to a GPT Image model, which changes the output and the cost — omit them unless you need them.
  • Models: prefer "automatic", which picks a model that supports the requested options. Pin one (e.g. gemini_3_1_flash_image, gpt_image_2) only when needed. Unsupported option/model combinations return 400. Full model list and limits: Generate images with the AI Gateway.
  • Cost preview: add dry_run: true to the same request to get { dry_run: true, model, usage: { base44_credits } } without generating or charging.
With the Vercel AI SDK, use models.imageModel("automatic") with generateImage and pass the Base44 extensions under providerOptions: { base44: { ... } } (the key is the name you gave createOpenAICompatible).

Generate videos

Video generation is an asynchronous job through the openai SDK’s videos client. videos.create() returns a job id right away; the video is ready later.
Poll from the frontend (or a later scheduled run), not inside one function invocation — a backend function is a bounded HTTP call:
  • Statuses: queued, in_progress, completed, failed. Treat queued/in_progress or hitting your polling limit as unfinished, not failed. For background flows, persist video.id (e.g. in an entity) and retrieve it in a later run.
  • Cost differs a lot by model. A 4-second clip without audio ranged from 48 credits (veo_3_1_lite) to over 350 (kling_3) when this was written. Preview with dry_run before choosing a model, and default to veo_3_1_lite unless the app needs another.
  • Models: there is no automatic — pass a specific model: veo_3_1_lite, veo_3_1_fast, seedance_2, seedance_2_5, seedance_2_fast, seedance_2_mini, kling_3, minimax_h3, minimax_h3_max, grok_imagine_video, grok_imagine_video_1_5. Supported seconds, resolution, aspect_ratio, and references differ per model, and some models reject generate_audio; unsupported values return 400, and an unknown model (including automatic) returns 404.
  • Request fields: model, prompt, seconds, aspect_ratio, resolution, generate_audio, seed, dry_run, and either frame_images (up to 2, each { type: "image_url", image_url: { url }, frame_type: "first_frame" | "last_frame" }) or input_references (up to 8 image/video/audio URLs, each { type: "image_url", image_url: { url } } and likewise for video_url / audio_url) — don’t combine the two. Reference URLs must be public HTTPS URLs.
  • Cost preview: videos.create({ ...request, dry_run: true }) returns HTTP 200 with usage.base44_credits and no job id — don’t poll it, and don’t return it as a 202.

Generate speech

  • Browser-only read-aloud, without stored audio or voice/style requirements: use the browser’s TTS API. This option does not apply to native apps.
  • Straightforward stored MP3 with a Core voice (river default, honey, sunny, storm, spark), optional ISO-639-1 language_code, up to 5,000 characters: use integrations.Core.GenerateSpeech({ text, voice?, language_code? }) → { url }.
  • Model choice, delivery instructions, speed, other formats, or voices outside Core: use the gateway. Follow this section when editing existing gateway speech code too.
POST /audio/speech returns completed audio bytes. It does not store a file or return a URL or job ID; there is no polling step. Call it from a backend function, following the shared rules and functions.md. text is the supplied spoken text; base44 is the function’s request client:
This example stores private MP3 audio. To store public audio instead, base44.asServiceRole.integrations.Core.UploadFile({ file }) returns file_url (see UploadFile). For private file_uri, call base44.integrations.Core.CreateFileSignedUrl({ file_uri }) and use signed_url for playback (see the signed-URL flow). Match the filename and content type when choosing another encoding. Prefer MP3 or WAV for playback; raw PCM requires decoding and cannot be played directly as an audio URL.
  • Web: render <audio controls src={signed_url}>; do not assume autoplay is allowed.
  • Native: use an installed React Native-compatible player or Linking.openURL(signed_url) to open saved audio externally. The native template has no audio-player package. Do not use HTML <audio>, new Audio, or speechSynthesis.
Cost preview: run this instead of the generation/upload block. dry_run: true returns JSON, without generating or charging; the estimate can differ from the final charge and contains no audio or file to play:
Request fields: Models and controls: omitted model or "automatic" resolves to gpt_4o_mini_tts.
  • OpenAI legacy voices: alloy, ash, coral, echo, fable, onyx, nova, sage, shimmer. OpenAI extended adds ballad, verse, marin, cedar.
  • ElevenLabs voices: honey, river, spark, storm, sunny.
  • Gemini voices: Kore, Puck, Zephyr, Charon, Leda, Fenrir (case-sensitive).
  • OpenAI input limit: 4,096 characters for all three models.
  • eleven_flash_v2_5 language codes: ar, bg, cs, da, de, el, en, es, fi, fil, fr, hi, hr, id, it, ja, ko, ms, nl, pl, pt, ro, ru, sk, sv, ta, tr, uk, zh.
Model discovery: GET /audio/speech/models returns { items, next_cursor }; GET /audio/speech/models/{model} returns one entry with its capabilities. Use these gateway catalogs for voices, formats and their defaults, speed ranges, instruction and language support, and input-length limits. The automatic entry’s resolves_to identifies the fixed default. Do not assume upstream provider features are exposed by the gateway.

AI decisions (Jev)

Structured evaluations: give the jev model some state and a set of questions, and get each answer back with probabilities. Use it to classify, score, or route; then apply your own thresholds and actions in deterministic code.
  • Needs ai 7.0.105+ — older versions (including the 7.0.16 code-agent pin above) don’t export experimental_evaluate.
  • The model id is exactly "jev". The package README’s "jev-latest" is rejected by the gateway.
Request:
  • state — string, JSON object, or array (an array is one state).
  • questions — non-empty; questions can’t depend on other answers in the same call.
    • choice: criteria required, an object of 1–255 named options (descriptions may be null).
    • score: criteria required, an array of 2–10 ordered rubric levels.
    • boolean: criteria optional, an object with only true/false keys (descriptions may be null).
    • instructions and criteria descriptions may be a string, JSON object, or JSON array.
Response (answer keys match question ids):

Models

  • Chat: any model available through InvokeLLM; automatic by default (see the code-agent rules above).
  • Images: automatic or a pinned image model — see Generate and edit images.
  • Videos: always a specific model — see Generate videos.
  • Speech: automatic is a fixed default; use the speech catalog for model-specific controls — see Generate speech.
  • Evaluations: jev.

Notes

  • Token: the service-role token for base44.asServiceRole.aiGateway; the caller’s token (or an empty string when unauthenticated) for base44.aiGateway.
  • Billing: metered per call against your app’s credit quota (same as InvokeLLM). If the app is out of credits, calls are rejected before the model runs.

Type Definitions