Skip to content

Add 2026-06-llmops-quickstart blog code - #91

Open
CEDipEngineering wants to merge 9 commits into
databricks-solutions:mainfrom
CEDipEngineering:add-2026-06-llmops-quickstart
Open

Add 2026-06-llmops-quickstart blog code#91
CEDipEngineering wants to merge 9 commits into
databricks-solutions:mainfrom
CEDipEngineering:add-2026-06-llmops-quickstart

Conversation

@CEDipEngineering

@CEDipEngineering CEDipEngineering commented Jun 5, 2026

Copy link
Copy Markdown

What this is

End-to-end LLMOps quickstart on Databricks — a customer support ticket classifier that carries an LLM agent through the full lifecycle: data ingestion → agent build → evaluation → approval → deployment → inference.

Accompanies an upcoming Databricks Community blog post (JIRA TLC-1077).

Architecture (2026 pattern)

  • Agent served as a Databricks App — a FastAPI agent server (mlflow.genai.agent_server @invoke handler), not a Model Serving endpoint. Serving agents as apps is the current recommendation; agents.deploy() is reserved for special custom cases.
  • LLM via a Unity AI Gateway model service — the agent calls the model service by its fully-qualified name (<catalog>.default.claude-sonnet-5, also gpt-oss-120b) through the AI Gateway. The model service is a Unity Catalog securable, so access control, rate limits, and payload logging live in UC. One env var (LLM_MODEL) switches models.
  • MLflow 3 GenAI evaluation gates the deploymlflow.genai.evaluate with a deterministic exact_match gate scorer plus the out-of-the-box Correctness judge; every prediction is captured as an MLflow Trace. agent-evaluate exits non-zero below the threshold (CI gate). A human approves before the app is deployed.

Contents

  • New folder 2026-06-llmops-quickstart/ (YYYY-MM-[name] convention)
  • agent_server/ (agent, evaluation, server), app.yaml, pyproject.toml + uv.lock
  • A Declarative Automation Bundle (databricks.yml) with the app resource, MLflow experiment, and the data-ingestion job (dev/prod targets)
  • README.md with prerequisite-knowledge, setup, governance recap, and a "before you call it done" checklist
  • LICENSE.md (unchanged) and CODEOWNERS

Testing

  • Agent server built and run locally; /invocations returns correct categories.
  • Model services confirmed on two workspaces (AWS + Azure) through the AI Gateway.
  • mlflow.genai.evaluate run: 90% exact-match, gate passed, traces logged.
  • Bundle validates and deploys (app resource + experiment + ingestion job).
  • Note: the Databricks Apps build installs from the workspace package proxy, which can return transient package-download errors; retry the deploy if a build fails on one.

Guidelines checklist

  • I have read the contribution guidelines
  • No sensitive information / no secrets
  • No PII; no external dataset (30 small, hand-written synthetic tickets)
  • Licenses listed in the README; Databricks license unchanged
  • Folder renamed to YYYY-MM-[folder_name]
  • Added myself to CODEOWNERS
  • Domain SME has reviewed the code and approved in a PR comment (pending)

Blog post link will be added to the README once published.

This pull request and its description were written by Isaac.

End-to-end LLMOps quickstart on Databricks (customer support ticket
classifier): MLflow ChatAgent on a Foundation Model API endpoint, an
evaluation gate that promotes a Champion in Unity Catalog, and a
Databricks Asset Bundle that deploys schema, experiment, and jobs with
batch + real-time inference.

- Folder follows the YYYY-MM-[name] convention
- README documents setup, structure, data (30 synthetic tickets, no PII), and licenses
- LICENSE.md is an unmodified copy of the repo Databricks license
- CODEOWNERS entry added for the new folder

I have read the contribution guidelines. No sensitive info, no PII, no
external dataset. Pending: SME code review approval and internal approval.

Co-authored-by: Isaac
CEDipEngineering and others added 5 commits June 5, 2026 16:13
…al, AI Gateway logging

Bring the quickstart current (2026) and align it with the MLOps quickstart's
step-by-step, best-practices structure. Tested end-to-end across three workspaces
(deploy → 4 jobs → approval → governed endpoint → batch + realtime).

- Evaluation: replace the hand-rolled accuracy loop with mlflow.genai.evaluate()
  using a deterministic exact_match gate scorer plus the built-in Correctness LLM
  judge; every row is captured as an MLflow Trace.
- Lifecycle: passing versions register as Challenger; new model_approval.py promotes
  Challenger→Champion only when run with --params approved=true, wired as a
  predecessor task to deployment.
- Governance: after agents.deploy(), enable AI Gateway inference-table payload logging
  on the agent endpoint (idempotent). Guardrails/rate limits documented as an
  FM-endpoint pattern (not supported on custom agent endpoints).
- Agent: normalize response content so reasoning models (Claude Sonnet 5, GPT-5) that
  return structured content blocks work, not just plain-string responses.
- Default LLM endpoint → databricks-claude-sonnet-5; pin mlflow>=3.4.0 in the logged model.
- Harden stale-deployment cleanup; README updated with the new flow, approval gate,
  governance notes, UC-privilege prerequisite, and a "before you call it done" checklist.
Per SME review, move off Model Serving / agents.deploy() to the current recommended
pattern: serve the agent as a Databricks App and call the LLM through a Unity AI
Gateway model service. Validated locally (agent server serves correctly) and against
real model services on two workspaces; MLflow 3 GenAI eval passes at 90%.

- agent_server/: FastAPI app (mlflow.genai.agent_server @invoke handler) classifying a
  ticket into one of five categories; calls the model service by fully-qualified name
  via the AI Gateway; reasoning-model-safe content handling.
- LLM via UAIG model service (LLM_MODEL = <cat>.default.claude-sonnet-5; also
  gpt-oss-120b); governance lives in Unity Catalog.
- Evaluation gates the app deploy: mlflow.genai.evaluate with an exact_match gate
  scorer + the out-of-the-box Correctness judge; agent-evaluate exits non-zero below
  threshold. Dropped UC model registration and Champion/Challenger aliases.
- Bundle: app resource + experiment + data-ingestion job (writes the support_tickets
  UC managed table). uv.lock via --exclude-newer for proxy-mirrored versions.
- Current naming throughout: Declarative Automation Bundles, UAIG, UC managed table,
  out-of-the-box. README adds prerequisite-knowledge, qs_catalog naming, training
  pointers, UC-privilege note, and a "before you call it done" checklist.
Link the free DevOps Essentials for Data Engineering (bundles/CI-CD) and Building
Agentic Applications on Databricks (agents, MLflow, evaluation) courses, per SME
suggestion to point readers at relevant training.

@kwulffert23 kwulffert23 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hi @CEDipEngineering , I left a couple of comments. Please let me know if you have questions.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@CEDipEngineering, not sure if you want to keep or rework the diagram to align with the code, as it's showing the agent deploy as model instead over apps.

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Good catch, updated the diagram.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@CEDipEngineering, same as above, I'd suggest reworking the diagram to align with the code, as it's showing the agent as mlflow model instead of agent on apps.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Both diagrams don't seem to be mentioned in the readme...

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Updated both diagrams, and added them to readme.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@CEDipEngineering the current .env.example is missing CATALOG_NAME and SCHEMA_NAME and LLM_MODEL

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Right, forgot to update the .env, should be fixed now.

Comment thread 2026-06-llmops-quickstart/README.md Outdated

```bash
databricks bundle deploy \
-v catalog_name=qs_catalog \

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@CEDipEngineering I might be outdated with the doc but isn't it --var?

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes, it is, hehe. Fixed here and on the Blog.

Comment thread 2026-06-llmops-quickstart/README.md Outdated
w = WorkspaceClient()
resp = w.api_client.do(
"POST",
f"/api/2.0/apps/{app_url}/invocations",

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@CEDipEngineering could you double check the invocation code, did it work? I had in mind /api/2.0/apps/{app_name}/invocations from https://github.com/databricks/databricks-sdk-py/blob/main/databricks/sdk/service/apps.py (line 3398)

@CEDipEngineering CEDipEngineering Aug 5, 2026

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

/api/2.0/apps/{app_name}/invocations does not work. I mixed it up with {app_url}/invocations, which is the endpoint I built on the app.

resources:
apps:
ticket_classifier:
name: "llmops-quickstart-classifier"

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

do we need a prefix in the app's name? e.g. if we deploy in the same workspace without any prefix the same name will be hardcoded in both targets dev and prod ending in resource already exists if running on the same workspace. It could be something like name: "llmops-quickstart-classifier-${bundle.target}"

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Agreed, done.

variables:
environment: dev

prod:

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@CEDipEngineering Shall we add root_path? Setting root_path pins prod to a single location regardless of who runs the deploy. https://github.com/databricks/devhub/blob/main/.agents/skills/databricks-core/declarative-automation-bundles.md

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I added it, however I decided to stick with the same approach as the MLOps repo I'm basing this on, which does root_path = /Workspace/Users/${workspace.current_user.userName}/.bundle/${bundle.name}/${bundle.target}
by default. I also added a comment there to explain why users might want to change this to something static, like /Workspace/Bundles/<project_name>/<bundle_name>/ and why using /Workspace/Shared is not a good idea.

CEDipEngineering added 3 commits August 5, 2026 16:39
… naming, prod root_path, diagrams

- Replace invalid `-v key=value` with `--var="key=value"` (the -v shorthand
  does not exist; the documented commands failed outright)
- Fix the inference snippet: a Databricks App is called on its own URL, not via
  a /api/2.0/apps/.../invocations control-plane endpoint (no such API)
- Add the missing LLM_MODEL, CATALOG_NAME, SCHEMA_NAME (plus optional
  ACCURACY_THRESHOLD, DATABRICKS_WAREHOUSE_ID) to .env.example
- Suffix the app name with ${bundle.target} so dev and prod can coexist in one
  workspace (development mode does not prefix app names)
- Pin the prod bundle root_path to a fixed identity instead of the deployer's
  home folder, keeping it out of world-writable /Workspace/Shared
- Regenerate both diagrams to match the App + UAIG model service architecture
  and reference them from the README; add the Mermaid sources
- Read the app source path from `bundle summary` rather than hardcoding it

Co-authored-by: Isaac
…he code

- app.yaml set LLM_MODEL to a bare endpoint name rather than a model-service
  fully-qualified name, contradicting the rest of the quickstart; note that the
  bundle's config.env overrides app.yaml when deploying through the bundle
- evaluate_agent.py required no CATALOG_NAME/SCHEMA_NAME and silently fell back
  to main.llmops_quickstart, so evaluating a prod deploy would score the dev
  schema and report a green gate against the wrong data
- evaluate_agent.py now fails with a clear message when MLFLOW_EXPERIMENT_ID is
  unset, rather than a RESOURCE_DOES_NOT_EXIST stack trace, and resolves .env
  relative to the package instead of the current directory
- agent.py docstring claimed an OpenAI client is used; it uses the SDK's HTTP
  client
- databricks.yml: prod_service_principal needs a default, since a variable with
  no default is required by every target and broke dev validation

Verified: dev and prod targets validate, and the evaluation runs end to end at
90.0% exact-match against the 30-ticket set.

Co-authored-by: Isaac
Sets root_path explicitly on dev and prod with ${workspace.current_user.userName},
the same pattern the companion MLOps Quickstart uses, so the two posts stay
consistent. Drops the prod_service_principal variable, so both targets validate
with no flags.

Leaves a comment recommending a static folder with controlled access, recreated
the same way on every workspace (e.g. /Workspace/Bundles/<bundle>/<target>), as
the stronger option for a real prod deployment.

Co-authored-by: Isaac
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants