GenAiHub

Last checked 8 October 2026 — responded normally.

Evaluate AI support agent quality and regressions with Google Sheets, Groq and Gmail

Quick overview

Live
Open / InstallLast updated October 8, 2026

Description

Quick overview This workflow runs a customer-support AI agent against Google Sheets test cases, uses Groq-hosted LLMs to evaluate response quality, logs results back to Google Sheets, compares run metrics to an approved baseline, and sends regression or review alerts via Gmail. How it works Starts when you manually execute the workflow. Reads test cases from Google Sheets, filters to only Active scenarios, and iterates through them in batches. Builds a customer-support prompt for each test case and generates an agent response using Google Gemini. Sends the question, expected facts/outcome, and the agent response to a Groq LLM to score relevance, reference accuracy, completeness, instruction compliance, and hallucination risk as JSON. Parses the evaluator JSON and appends per-test results (including the agent response and overall result) to an Evaluation_Results sheet in Google Sheets. Aggregates the current run’s results, selects the latest previously reviewed baseline from Google Sheets, and computes pass-rate regression plus metric deltas. Writes the run-level metrics back to Google Sheets and routes the outcome to either mark the run successful or send a Gmail warning/alert to the QA team. Setup Create a Google Sheets file with Test_Cases, Evaluation_Results, and Baseline_Metrics tabs (including columns referenced in the workflow like Active, Test_ID, Question, Expected_Facts, Expected_Outcome, and Reviewed). Add Google Sheets credentials in n8n and replace YOUR_GOOGLE_SHEET_ID in all Google Sheets nodes with your spreadsheet ID. Add credentials for Google Gemini and Groq (used for the agent and evaluator chat models) and select the models you want to run. Add Gmail credentials and update the recipient address (qa-team@example.com) and any email content/subjects to match your QA process. Ensure at least one baseline row in Baseline_Metrics is marked Reviewed=true so the workflow can compare the current run against an approved baseline. An n8n automation workflow template by Fahim Jilani.

Community Metrics

Views
0
Avg Rating
N/A
Ratings
0
Likes
0
Comments
—

Author

Fahim Jilani

Platform

web

Pricing model

paid

Categories

  • Automation
  • AI

Tags

  • n8n
  • workflow
  • google-sheets
  • gmail
  • code
  • ai-agent
  • basic-llm-chain
  • google-gemini-chat-model
  • groq-chat-model

Capabilities

  • 7 nodes

Information

TypeWorkflow
Sourcen8n
AddedSeptember 29, 2026

Comments

?

Others in the same category, ranked by how often they are opened.