How Prisma Built a Software Factory With a Slack Coding Agent and an AI SRE

Gremlin turns Slack conversations into pull requests; Sherlog investigates telemetry and runbooks before an engineer starts digging.

500K+

Developers use Prisma

~2 min

From Sentry stack trace to PR

Prisma builds the data and runtime stack for TypeScript applications. Its products are used by more than 500k developers.

Built with Mastra, Gremlin and Sherlog are early parts of Prisma's software factory that improves how the company ships and operates the platform. Gremlin turns Slack discussions into PRs. Sherlog checks telemetry, incidents, and runbooks before the on-call engineer starts digging. It already reduces investigation time from hours to minutes.

Turning Slack discussions into code

When you ship software, lots of minor points of friction get created. They accumulate in a backlog because they are too small for a single PR. Or when you get to them while working on a larger project, they slow down the release process.

The context for those fixes often already existed in a Slack thread, so Prisma built Gremlin, a cloud agent that picks up that conversation and ships fixes in minutes.

It has access to GitHub and Prisma's knowledge base, and devs can chat with it in Slack or Mastra Studio. It's built on Prisma Compute and Prisma Postgres. Other agents can also call Gremlin through the API.

A Gremlin task starts when a user tags it in a discussion, attaches a brief, or asks it to summarize the thread and create one.

Prisma's first design sent the brief directly to a coding sandbox, but it couldn't recover from a failed run or push back when the request was incomplete. The team placed a Mastra agent between Slack and the sandbox.

"I decided to build on Mastra because I didn't want to spend time building a harness and a workflow. That's not the problem we're trying to solve at Prisma. So I chose a tool that's open source, very modular, clear separation of concerns, and a wild wealth of primitives. Mastra just made it incredibly easy to build complex systems." — Tyler Hogarth, Head of Engineering

Gremlin now asks clarifying questions, reviews the result, and decides whether the coding agent needs another pass. It then confirms the request and provisions an isolated coding sandbox.

The Mastra agent also acts as a security layer. It holds the credentials, mints fresh tokens, injects the required environment variables, and clones the repository into the sandbox. This keeps the sandbox's access narrow and temporary.

The coding agent makes the change and returns the result to Gremlin. Gremlin can inspect the work and open a PR for the team.

In an early test, Gremlin turned a Sentry stack trace into a PR in about two minutes. The agent now handles small bugs and web changes, and can even build larger features.

Investigating production incidents faster

In the past, when an incident occurred, an on-call engineer had to search several systems, collect telemetry, check runbooks, and test possible causes. This could take hours, often in the middle of the night.

Prisma built Sherlog, a Slack-native SRE assistant, to shorten this time to minutes. Engineers can message Sherlog, tag it in a channel, or let it watch the production-incidents channel, and the agent does the first investigation before a person starts digging.

Sherlog uses ClickHouse for application metrics, Axiom for traces and logs, Grafana for incident management, and Prisma's Ignite repository for runbooks and product context. It can publish the result as a report in Prisma Storage.

The main challenge when building this was that Axiom queries produced too much noisy data. Sherlog's context filled up quickly with failed or unbounded queries, and the agent was losing important incident info.

As a fix, the team added a step: Sherlog now sends each Axiom hypothesis to a separate query-executor agent. That agent starts with a clean context, investigates one question, and returns only the useful signal. It can run several of these investigations in parallel.

Sherlog originally ran in n8n, but Prisma didn't want to keep editing it through a visual workflow UI. So they moved to Mastra, where they can keep the system in TypeScript and ask a coding agent to change the code if needed.

The agent already makes engineers' lives easier. For one performance report, Sherlog explored Prisma's ClickHouse cluster. It found the services contributing the most CPU and memory usage, analyzed errors, and charted week-over-week changes. It took the agent about 10 minutes. Doing it manually would take 1–2 hours.

Building the broader software factory

Gremlin and Sherlog are two parts of a larger internal agent system: a software factory. Prisma's first Mastra agent was a simple Changelog bot. They have since added Gizmo for code review and specialist agents such as Sherlog's Axiom query executor.

Prisma hosts the agents in a single Mastra deployment, where they can share tools but stay isolated. Other agents can call Gremlin through its API. Sherlog delegates telemetry research to the Axiom executor. Gremlin sends self-improvement changes through an external review agent and then to a person for approval.

Prisma runs the system on open-weight models and assigns models by role.

Gremlin and Sherlog use the reasoning role. Gizmo uses code review. The Axiom executor uses data analysis, and Changelog uses summarization. Prisma also defines roles for classification and structured extraction tasks.

The mappings live in one configuration file with fallbacks. When Prisma wants to test a new model, it can change the primary model for a role in one place, and every agent using that role picks up the change.

Expanding Gremlin and Sherlog

Prisma is building toward that larger system one task at a time, adding another agent where the pattern proves useful.

Sherlog's next goal is better first-line incident triage without unsafe access or overloaded context. Gremlin is moving toward deeper integration with Prisma's open-source repository.

Start building today

Quickstart