Skip to content

Display & model ​

Configure which model runs triage and how results appear on Pulse detail. In the Exhale app this is AI → Display & model. Workspace admins can save; members and viewers can read the current values.

Related: Prompt & behavior · Quality & feedback


Model & provider ​

Which LLM runs triage for this workspace: the active model, an optional customer API key, monthly run ceilings, and a backup model when the primary provider errors.

Open this group when you are choosing a model, attaching a Business key, or checking monthly run caps — not for how summaries look on Pulse detail (that is Presentation).

Cards inside this group vary by plan: active model is Starter and above; Anthropic and xAI BYOK are Business or higher; fallback model is Growth or higher; inference limits apply on all plans.

Active model ​

The model that powers each triage run for this organization. Paid plans can pick from the allowed list; Trial uses the platform default.

Change it if you want a different quality, latency, or cost profile. New runs pick up the saved model; already-triaged Pulses keep their stored result until you retry.

Self-serve model selection requires Starter or higher.

Bring your own API key ​

Use your Anthropic API key for Claude triage instead of platform inference. Keys are encrypted at rest and never returned in API responses. Test verifies the key without exposing it.

Use this when Business policy requires inference under your Anthropic account. Clear the key to return to platform inference.

Business or higher. Lower tiers see a locked card.

Bring your own xAI API key ​

Use your xAI API key for Grok triage instead of platform inference. Keys are encrypted at rest and never returned. Claude remains available via the platform Anthropic key or Anthropic BYOK.

Use this when Business policy requires Grok inference under your xAI account. Selecting a Claude model does not use this key.

Business or higher. Growth can still select Grok when Exhale provides a platform xAI key.

Inference limits ​

Monthly triage-run and concurrent-call ceilings for this workspace. You can lower them within the plan maximum. These are plan ceilings, not a live token-by-token meter. The badge shows monthly runs used versus your cap.

Lower the caps if you want a safety brake. Plan maxima are shown on the card.

Available on all plans. Maxima scale with the workspace tier.

Fallback model ​

A secondary model used once when the primary provider returns a rate limit or error. It is not used for ordinary runs.

Set a fallback if you would rather get a slower or cheaper result than fail triage when the primary model is unavailable. Choose None to disable.

Growth or higher. Trial and Starter see a locked card.

Presentation ​

How triage reasoning, confidence, summary length, and language appear to operators. These settings do not change which model runs.

Adjust these when Pulse detail feels too terse, too noisy, or when the team wants summaries in a language other than English.

Explainability, summary detail, and confidence display are on all plans. Output language requires Growth or higher.

Display ​

When enabled, Pulse detail can show which payload signals informed the triage result (explainability hints), if the run produced them.

Turn it on while you are tuning prompts or teaching the team how triage works. Turn it off if the extra detail is noise during a page.

Available on all plans.

Summary detail ​

How verbose new triage summaries are on Pulse detail and in notifications. Concise is shorter; Expanded includes more supporting detail.

Switch to Expanded when on-call wants more context in the first read. Concise is better for noisy queues. Applies to newly generated runs, not stored results.

Available on all plans.

Confidence display ​

Whether a confidence line appears on Pulse detail triage results (Triage → Results). Independent of explainability hints. List cards do not show confidence.

Hide pills if the team finds confidence scores distracting. Showing them helps spot low-confidence runs to review.

Available on all plans.

Output language ​

Language for newly generated triage summaries and other operator-facing LLM text. Does not translate stored results or incoming alert titles.

Independent of UI language. The Exhale interface language is Settings → Preferences (all plans). Triage summary language is this Output language control (Growth or higher). Changing one does not change the other. Slack and Teams channels are shared. Trial and Starter stay on English for model output. The model can use languages (for example pt / ja) that the interface does not ship yet.

This Output language is also the default when generating a Postmortem, Business impact, or FAQ. Operators can pick a different language on that report’s empty state; the choice is stored on the report and does not write back to this Display setting. Refresh and Enhance reuse the stored report language. Pulse titles and stored triage text stay as authored.

JSON keys stay English (summary, recommended_next_steps). Only string values follow output language.

Set this when the on-call team wants triage JSON in a language other than English.

Growth or higher. Trial and Starter stay on English.

Exhale by Kolstrom Systems LLC