Aptex AI logo Aptex AIBack to Aptex AI
Developer resources

Aptex AI API Documentation

This page documents the current server-side generation endpoint used by the Aptex AI website. It is an internal application endpoint, not a stable public developer SDK or a promise of long-term API compatibility.

Current API status

The current application exposes a Netlify Function at POST /.netlify/functions/generate-reply. The browser sends the request to Aptex AI's server-side function; the function authenticates with Groq using the server-side GROQ_API_KEY environment variable and returns a JSON response. The caller does not provide the Groq key.

Text generation

Text flow: text input → Aptex AI server function → Groq openai/gpt-oss-20b → generated reply.

The request body is JSON. The main text fields currently used by the application are message, context, situation, category, vibe, tone, personality, format, modifier, primaryIntent, inputIntent, and ideaMode.

{
  "message": "you really thought I was gonna text you first 😂",
  "context": "",
  "category": "reply",
  "vibe": "Funny",
  "tone": "Chill",
  "personality": "Auto",
  "format": "One-liner",
  "inputMode": "text"
}

Successful text generation returns JSON in the form {"text":"..."}. The response contains the generated reply only.

Screenshot generation

Screenshot flow: screenshot → vision processing → conversation understanding → selected vibe/category → generated reply.

Set inputMode to "screenshot" and provide the screenshot as a data URL in imageData. The server sends the image to Groq's configured vision model qwen/qwen3.8-27b. The vision prompt instructs the model to identify the relevant incoming message, use nearby visible context, ignore unrelated interface elements, and return only the send-ready reply.

{
  "message": "",
  "inputMode": "screenshot",
  "imageData": "data:image/png;base64,<base64 image data>",
  "category": "flirty",
  "vibe": "Flirty",
  "tone": "Chill",
  "personality": "Auto",
  "format": "One-liner"
}

Screenshot limits and image handling

The current app accepts image/png, image/jpeg/image/jpg, and image/webp. One screenshot is processed per generation request.

The website UI rejects an original image larger than 12 MB. Before the request is sent, the browser may resize images above 2200 px on their longest dimension and may convert oversized or resized images to JPEG at approximately 0.82 quality. The server enforces a maximum decoded screenshot size of 3,500,000 bytes and a maximum overall request body size of 5,000,000 bytes. The browser also rejects a compressed screenshot data URL above 4.75 MB.

No screenshot file is uploaded to a separate Aptex storage system by this endpoint. The application keeps the screenshot in browser memory for the active workflow, sends it to the Netlify Function, and the function processes it in request memory before sending it to Groq.

Language matching

Aptex AI attempts to respond in the language of the supplied conversation. In Text mode, language is inferred from the submitted message. In Screenshot mode, language is inferred from the visible conversation. Mixed-language conversations are handled by following the dominant language naturally while preserving common conversational terms where appropriate. This is an inference behavior, not a guarantee of perfect language detection.

Category and vibe handling

The selected category/vibe is passed into the server-side generation prompt and affects the reply's personality. The existing application categories and related controls remain available; unsupported category IDs are normalized to the server's safe fallback category rather than creating a new behavior.

Authentication and environment

The current function does not require a caller-supplied API key. The Groq credential remains server-side as GROQ_API_KEY. Text generation uses GROQ_MODEL (with the application default openai/gpt-oss-20b), while Screenshot mode uses GROQ_VISION_MODEL and the configured value qwen/qwen3.8-27b. These environment variables must not be exposed to the browser.

Request limits

The server limits a message to 6,000 characters, supplied context to 2,500 characters, and the overall JSON request body to 5,000,000 bytes. A lightweight per-instance guard allows up to 6 requests per minute for an observed client key. Because Netlify can run multiple function instances, that guard is not a guaranteed global quota.

Error responses

The function returns JSON with a specific code and user-safe error message for many failure cases. Important current screenshot errors include:

SCREENSHOT_MISSING
SCREENSHOT_UNSUPPORTED
SCREENSHOT_TOO_LARGE
SCREENSHOT_INVALID
SCREENSHOT_UNREADABLE
NO_CONVERSATION_DETECTED
AI_VISION_NOT_CONFIGURED
AI_VISION_MODEL_UNAVAILABLE
AI_VISION_AUTH_ERROR
AI_VISION_RATE_LIMIT
AI_VISION_TIMEOUT
AI_VISION_BAD_REQUEST
AI_VISION_INVALID_RESPONSE
AI_VISION_PROVIDER_ERROR

Text generation uses analogous configuration, authentication, model, rate-limit, timeout, malformed-response, and provider-error codes. API responses never include the Groq API key or internal stack traces.

Contact

For API questions, contact vibegenaii@gmail.com.