usefulapi

Together AI MCP server. Run Together AI chat, embeddings and images; manage fine-tunes, batches and endpoints.

Claude

  1. Open Settings → Connectors → Add custom connector
  2. Paste this URL:
    https://together-ai.usefulapi.io/mcp
  3. Authenticate with Together AI when prompted

Cursor · VS Code · Windsurf · Cline

Add to your MCP config, then reload & authorize:

{
  "mcpServers": {
    "together-ai": {
      "url": "https://together-ai.usefulapi.io/mcp"
    }
  }
}
live25 toolsFree 100 tool calls / monthPro $9/mo · $90/yr

Tools 25

ToolTypeWhat it does
together_whoamiread
Who am I
Identify the API key: its organization, project and project slug. The project slug forms the `<project_slug>/<endpoint_slug>` model name for dedicated-endpoint inference. A cheap way to confirm the key works. Together: GET /whoami.
together_list_modelsread
List models
List Together's models with type (chat, language, code, image, embedding, moderation, rerank), context length, organization, license and per-token pricing. Together: GET /models.
together_list_filesread
List files
List uploaded data files (fine-tune, eval and batch-api inputs, plus job outputs) with size, type, purpose and validation status. Together: GET /files.
together_get_fileread
Get one file
Fetch one file's metadata, including its processing_status and validation_report (why a fine-tune training file was rejected). Together: GET /files/{id}.
together_list_fine_tunesread
List fine-tuning jobs
List fine-tuning jobs with status, base model, output model name and training settings. Together: GET /fine-tunes.
together_get_fine_tuneread
Get one fine-tuning job
Fetch one fine-tuning job: status, progress, hyperparameters, token counts, cost and the output model name. Together: GET /fine-tunes/{id}.
together_list_fine_tune_eventsread
List a fine-tuning job's events
List the event log of one fine-tuning job (queued, started, checkpoint saved, epoch completed, errors). The first place to look when a job failed. Together: GET /fine-tunes/{id}/events.
together_list_batchesread
List batch jobs
List batch inference jobs with status, progress, model and input/output/error file ids. Together: GET /batches.
together_get_batchread
Get one batch job
Fetch one batch job: status (VALIDATING, IN_PROGRESS, COMPLETED, FAILED, EXPIRED, CANCELLED), progress, and the output_file_id / error_file_id once done. Together: GET /batches/{id}.
together_list_endpointsread
List endpoints
List endpoints with model, owner and state (PENDING, STARTING, STARTED, STOPPING, STOPPED, ERROR). Use mine=true and type=dedicated to see what is running on your account and billing by the minute. Together: GET /endpoints.
together_get_endpointread
Get one endpoint
Fetch one dedicated endpoint: state, model, hardware, autoscaling bounds and display name. Together: GET /endpoints/{endpointId}.
together_list_hardwareread
List hardware
List hardware configurations for dedicated endpoints with GPU type/count/memory and price in cents per minute. Pass a model to get only compatible configurations with live availability. Together: GET /hardware.
together_list_evaluationsread
List evaluation jobs
List LLM-as-a-judge evaluation jobs (classify, score, compare) with status, parameters and results once completed. Together: GET /evaluation.
together_chat_completionwrite
Chat completion
Run a chat completion on a Together model (billed per token). Non-streaming. For a dedicated endpoint pass its `<project_slug>/<endpoint_slug>` as the model. Together: POST /chat/completions.
together_create_embeddingswrite
Create embeddings
Generate vector embeddings for one or more texts (billed per token). Together: POST /embeddings.
together_generate_imagewrite
Generate an image
Generate images from a prompt (billed per image/megapixel). Returns image URLs by default rather than base64, to keep responses small. Together: POST /images/generations.
together_create_fine_tunewrite
Create a fine-tuning job
Start a fine-tuning job on an uploaded training file (billed per token processed). Stop it with together_cancel_fine_tune. Together: POST /fine-tunes.
together_cancel_fine_tunewrite
Cancel a fine-tuning job
Cancel a running fine-tuning job. Cannot be resumed, but a new job can continue from its last checkpoint via from_checkpoint. Together: POST /fine-tunes/{id}/cancel.
together_create_batchwrite
Create a batch job
Start an asynchronous batch job over an uploaded JSONL input file (purpose batch-api), at a discount to real-time inference. Together: POST /batches.
together_cancel_batchwrite
Cancel a batch job
Cancel a batch job that has not finished. Together: POST /batches/{id}/cancel.
together_create_endpointwrite
Create a dedicated endpoint
Deploy a model on dedicated GPUs. The endpoint STARTS AUTOMATICALLY and bills per minute of uptime until stopped — set inactive_timeout to auto-stop it, and use together_list_hardware for valid hardware ids. Together: POST /endpoints.
together_start_endpointwrite
Start a dedicated endpoint
Start a stopped dedicated endpoint. It bills per minute of uptime until stopped. Reversible with together_stop_endpoint. Together: PATCH /endpoints/{endpointId} with state=STARTED.
together_stop_endpointwrite
Stop a dedicated endpoint
Stop a running dedicated endpoint, which stops its per-minute billing. Requests to it fail until it is started again. Together: PATCH /endpoints/{endpointId} with state=STOPPED.
together_ai_usage_statusmeta
Usage status (free-tier meter)
Report the caller's current free-tier usage this month: calls used, monthly limit, remaining, and whether the cap is reached. Read-only; does not count against the meter.
together_ai_upgrademeta
Upgrade to Pro (unlimited)
Subscribe to the Pro plan for UNLIMITED Together AI tool calls (the free tier caps monthly usage). Choose monthly ($9/month) or yearly ($90/year — 2 months free) billing. Returns a Stripe Checkout link to open in your browser; after payment your account upgrades automatically. Read-only; does not count against the meter.

Pricing

PlanPriceLimit
Free$0100 tool calls / month
Proper user$9/mo · $90/yrUnlimited

This is a Model Context Protocol endpoint — meant to be connected from an AI client, not opened in a browser. An invalid_token response at the URL is the auth gate working as designed; clients authenticate automatically.