AI Testing Repos

AI-powered testing & quality automation tools on GitHub · updated daily

Inspired by goodailist.com

⚙️392repos
⭐147.5Ktotal stars
📈0added this week
🏆Othertop category
392 repos
1 / 4
#Repo ↓Stars ↓1d7dForksDescriptionCategoryLanguageUpdated
1
vxcontrol
pentagi
17,608——2,408Fully autonomous AI Agents system capable of performing complex penetration testing tasksSecurity TestingGo2026-05-31
2
raga-ai-hub
RagaAI-Catalyst
16,170——3,600Python SDK for Agent AI Observability, Monitoring and Evaluation Framework. Includes features like agent, llm and tools tracing, debugging multi-agentic system, self-hosted dashboard and advanced analytics with timeline and execution graph view Performance TestingPython2026-02-11
3
cvat-ai
cvat
16,017——3,702Computer Vision Annotation Tool (CVAT) is a leading platform for building high-quality visual datasets for vision AI. It offers open-source, cloud, and enterprise products, as well as labeling services, for image, video, and 3D annotation with AI-assisted labeling, quality assurance, team collaboration, analytics, and developer APIs.Visual TestingPython2026-06-10
4
alibaba
MNN
14,384——2,218MNN is a blazing fast, lightweight deep learning framework, battle-tested by business-critical use cases in Alibaba. Full multimodal LLM Android App:[MNN-LLM-Android](./apps/Android/MnnLlmChat/README.md). MNN TaoAvatar Android - Local 3D Avatar Intelligence: apps/Android/Mnn3dAvatar/README.mdOtherC++2026-03-05
5
web-infra-dev
midscene
13,653——1,036AI-powered, vision-driven UI automation for every platform.E2E TestingTypeScript2026-06-10
6
GreyDGL
PentestGPT
13,627——2,360Automated Penetration Testing Agentic Framework Powered by Large Language ModelsSecurity TestingPython2026-06-07
7
evidentlyai
evidently
7,590——861Evidently is ​​an open-source ML and LLM observability framework. Evaluate, test, and monitor any AI-powered system or data pipeline. From tabular data to Gen AI. 100+ metrics.Test GenerationJupyter Notebook2026-05-02
8
Giskard-AI
giskard-oss
5,425——468🐢 Open-Source Evaluation & Testing library for LLM AgentsSecurity TestingPython2026-06-10
9
qodo-ai
qodo-cover
5,424——520Qodo-Cover: An AI-Powered Tool for Automated Test Generation and Code Coverage Enhancement! 💻🤖🧪🐞Test GenerationPython2026-04-05
10
GH05TCREW
pentestagent
2,623——519PentestAgent is an AI agent framework for black-box security testing, supporting bug bounty, red-team, and penetration testing workflows.Security TestingPython2026-05-27
11
Armur-Ai
Pentest-Swarm-AI
1,771——358Autonomous penetration testing using a swarm of AI agents. Orchestrates recon, classification, exploitation, and reporting specialists with ReAct reasoning — supports bug bounty, continuous monitoring, and CTF modes. Built with Go, Claude API, and 7+ native security tools.Security TestingGo2026-06-06
12
kubeshop
testkube
1,605——163☸️ The Open Testing Platform for AI-Driven Engineering TeamsObservabilityGo2026-06-10
13
zakirkun
guardian-cli
1,450——305Guardian is a production-ready AI-powered penetration testing automation CLI tool that leverages Google Gemini and LangChain to orchestrate intelligent, step-by-step penetration testing workflows while maintaining ethical hacking standards.Test GenerationPython2026-05-29
14
JoasASantos
NeuroSploit
1,122——273NeuroSploit is an advanced, AI-powered penetration testing framework designed to automate and augment various aspects of offensive security operations. Leveraging the capabilities of large language models (LLMs).Security TestingPython2026-03-29
15
test-zeus-ai
testzeus-hercules
1,040——162Hercules is the world’s first open-source testing agent, enabling UI, API, Security, Accessibility, and Visual validations – all without code or maintenance. Automate testing effortlessly and let Hercules handle the heavy lifting! ⚡E2E TestingPython2026-06-03
16
SanMuzZzZz
LuaN1aoAgent
1,015——158LuaN1aoAgent is a cognitive-driven AI hacker. It is a fully autonomous AI penetration testing agent, using dual-graph reasoning.Security TestingPython2026-04-13
17
bug0inc
passmark
950——165The open-source Playwright library for AI browser regression testing with intelligent caching, auto-healing, and multi-model verification.E2E TestingTypeScript2026-06-08
18
MGdaasLab
WHartTest
934——125WHartTest 是一款AI驱动的测试自动化平台,实现从需求到可执行测试用例的自动化生成与管理,帮助测试团队提升效率与覆盖率。 (WHartTest is an AI-driven test automation platform that automates the generation and management of executable test cases from requirements, helping testing teams improve efficiency and coverage.)E2E TestingPython2026-06-04
19
PurpleAILAB
Decepticon
932——175Autonomous Multi-Agent Based Red Team Testing Service / AI hackerTest GenerationPython2026-03-26
20
CyberSecurityUP
NeuroSploit
910——234NeuroSploit is an advanced, AI-powered penetration testing framework designed to automate and augment various aspects of offensive security operations. Leveraging the capabilities of large language models (LLMs).Security TestingPython2026-02-24
21
alumnium-hq
alumnium
906——92End-to-end testing with AIE2E TestingTypeScript2026-06-09
22
langwatch
scenario
895——65Agentic testing for agentic codebasesOtherPython2026-06-10
23
Agent-Field
SWE-AF
844——133Autonomous software engineering fleet of AI agents for production-grade PRs on AgentField: plan, code, test, and ship.Test GenerationPython2026-06-01
24
zakirkun
deep-eye
831——177An advanced AI-driven vulnerability scanner and penetration testing tool that integrates multiple AI providers (OpenAI, Grok, OLLAMA, Claude) with comprehensive security testing modules for automated bug hunting, intelligent payload generation, and professional reporting.Test GenerationPython2026-02-03
25
ScrapeGraphAI
scrapecraft
656——99🤖 AI-powered web scraping editor with visual workflow builder. Build, test & deploy web scrapers using natural language. Powered by ScrapeGraphAI & LangGraph.Visual TestingPython2025-12-26
26
pikpikcu
airecon
643——100AIRecon is an autonomous cybersecurity agent that combines a self-hosted Large Language Model (Ollama) with a Kali Linux Docker sandbox and a Textual TUI. It is designed to automate security assessments, penetration testing, and bug bounty reconnaissance — without any API keys or cloud dependency.Security TestingPython2026-04-23
27
CopilotKit
aimock
623——41Mock everything your AI app talks to — LLM APIs, MCP, A2A, AG-UI, vector DBs, search. One package, one port, zero dependencies.OtherTypeScript2026-06-10
28
trueleaf
apiflow
603——99A modern API workspace that works both online and offline — combining API documentation, testing, mock, and AI-powered automation in one lightweight toolOtherTypeScript2026-06-10
29
ServiceNow
AgentLab
586——116AgentLab: An open-source framework for developing, testing, and benchmarking web agents on diverse tasks, designed for scalability and reproducibility.Performance TestingPython2026-03-17
30
PacificAI
langtest
560——49Deliver safe & effective language modelsPerformance TestingPython2026-04-22
31
Pacific-AI-Corp
langtest
551——50Deliver safe & effective language modelsPerformance TestingPython2026-02-19
32
modal-labs
devlooper
471——31A program synthesis agent that autonomously fixes its output by running tests!OtherPython2024-09-19
33
RevoltDevScript
Revolt-Script
461——1Revolt is the #1 Edgenuity automation tool featuring Auto Quiz, Auto Essay, Auto Advance, AI-powered writing with humanization, Auto Vocabulary, Auto Journal, and more. A semi-AFK script that handles quizzes, tests, exams, projects, and essays with ease. Compatible with desktop and mobile. Visit https://revolt.ly to get started.Other—2026-04-16
34
project-codeguard
rules
410——56Project CodeGuard is an AI model-agnostic security framework and ruleset that embeds secure-by-default practices into AI coding workflows (generation and review). It ships core security rules, translators for popular coding agents, and validators to test rule compliance.Test GenerationPython2026-01-29
35
SHAdd0WTAka
Zen-Ai-Pentest
391——66🛡⚔️AI-Powered Penetration Testing Framework with automated vulnerability scanning, multi-agent system, and compliance reporting🛡⚔️Security TestingPython2026-06-08
36
Ktovoz
Skiritai
383——16AI-powered test automation framework that explores test paths with AI and generates replayable scripts for 30x faster execution.Test GenerationPython2026-05-30
37
z3n70
Frida-Script-Runner
363——75Web-based Frida framework and toolkit for Android & iOS penetration testing, mobile security, and dynamic analysis, featuring AI-assisted Frida script generation.Test GenerationJavaScript2026-01-26
38
faiscadev
fakecloud
316——19Free, open-source AWS emulator. LocalStack alternative: 40 services, 2,935 operations, true 100% Smithy conformance (99,678/99,678 variants pass). No account, no auth token, no paid tier.Integration TestingRust2026-06-10
39
CyberStrikeus
CyberStrike
316——58AI-powered offensive security agent with 7,300+ actionable security skills. Autonomous pentesting powered by MITRE ATT&CK (2,000+ Atomic tests), CIS Benchmarks (1,500+ controls), OWASP, NIST. Lazy-loading, zero context pollution. Your AI red team.Security TestingTypeScript2026-06-09
40
aielte-research
HackSynth
310——51LLM Agent and Evaluation Framework for Autonomous Penetration Testing Security TestingPython2025-06-24
41
ai-dashboad
flutter-skill
284——38AI-powered E2E testing for 10 platforms. 253 MCP tools. Zero config. Works with Claude, Cursor, Windsurf, Copilot. Test Flutter, React Native, iOS, Android, Web, Electron, Tauri, KMP, .NET MAUI — all from natural language.E2E TestingDart2026-05-21
42
pensarai
apex
283——49AI-powered offensive security testing using autonomous agents, directly in your terminal.Security TestingTypeScript2026-06-09
43
jd-opensource
JoySafeter
282——53🚀 JoySafeter: An enterprise AI Agent Platform—Not just chatting. building、running、testing, and tracing autonomous Agent Teams with visual orchestration...Visual TestingPython2026-05-29
44
capture0x
AdStrike
257——52AI-powered modular Active Directory red-team framework for authorized penetration testing, AD enumeration, attack-path analysis, Kerberos/ADCS workflows, reporting, operator automation, and MCP server integration.Security TestingPython2026-06-02
45
Coff0xc
AutoRedTeam-Orchestrator
247——50Enterprise AI Red Team Platform | 企业级AI红队平台 | 132 MCP Tools | Pure Python Engines | SDK+CLI+MCP | Auto-Download sqlmap/nuclei/ffuf | Production C2 | LLM Enhanced | Docker Sandbox | SARIF CI/CD | 1980 TestsSecurity TestingPython2026-05-18
46
gsd-build
gsd-browser
238——15A fast, native browser automation CLI built from the ground up for AI agents, powered by Chrome DevTools Protocol. 63 commands covering navigation, interaction, screenshots, accessibility, network mocking, visual diffing, test generation, and more — all from a single binary.E2E TestingRust2026-05-03
47
praetorian-inc
augustus
234——30LLM security testing framework for detecting prompt injection, jailbreaks, and adversarial attacks — 190+ probes, 28 providers, single Go binarySecurity TestingGo2026-06-10
48
Corpus-OS
corpusos
224——7Open-source protocol suite standardizing LLM, Vector, Graph, and Embedding infrastructure across LangChain, LlamaIndex, AutoGen, CrewAI, Semantic Kernel, and MCP. 3,330+ conformance tests. One protocol. Any framework. Any provider.OtherPython2026-05-08
49
LLAMATOR-Core
llamator
214——20Red Teaming python-framework for testing chatbots and GenAI systems.Security TestingPython2026-05-20
50
inhouseseo
superseo-skills
199——2711 Claude skills for SEO: page audits, linkbuilding, article writing, E-E-A-T audits, semantic gap analysis, link building. Methodology from Koray Tuğberk, Kyle Roof, and Lily Ray, plus a generation-time anti-AI-slop ruleset. Production-tested at InhouseSEOTest Generation—2026-05-12
51
Hellsender01
LLMMap
195——22Automated prompt injection testing framework for LLM-integrated applications with dual-LLM architecture.OtherPython2026-03-14
52
DanielSuo117
velocitai
183——0Next-generation UI automation harness powered by Python & Playwright. Chat with AI Agents to seamlessly generate, architect, and execute enterprise-grade UI test code.E2E TestingHTML2026-05-24
53
KHenryAegis
VulnBot
175——29The repository of VulnBot: Autonomous Penetration Testing for A Multi-Agent Collaborative Framework.OtherPython2025-04-07
54
CatchTheTornado
open-agents-builder
173——36AI Agents are missing the UI! We're here to change it. Build Business AI Agents for your company: business workflows, API's, bookings, e-commerce, social commerce, b2b, CPQ, intake forms, NPS tests, made-to-order use casesOtherTypeScript2025-12-08
55
AISquare-Studio
AISquare-Studio-QA
166——9Setting up QA testing agents using playwright and crewAIE2E TestingPython2026-04-17
56
fugazi
test-automation-skills-agents
161——27A practical library of agents, instructions, and skills designed specifically for QA Automation Engineers, focusing on production-oriented solutions.OtherJava2026-05-25
57
ShirleyRex
assertify.io
157——121AI-powered assistant that turns your project context into end-to-end test strategies, detailed scenarios, and ready-to-run boilerplate across popular frameworks.E2E TestingTypeScript2026-03-15
58
DebugBase
glance
151——23AI-powered browser automation MCP server for Claude Code. Navigate, click, screenshot, test — all from your terminal.E2E TestingTypeScript2026-04-16
59
pymc-labs
decision-lab
150——12Run tested, autonomous agent workflows on your data for meaningful decision-making OtherPython2026-06-03
60
fzn0x
watchtower
141——24Watchtower is a simple AI-powered penetration testing automation CLI tool that leverages LLMs and LangGraph to orchestrate agentic workflows that you can use to test your websites locally. Generate useful pentest reports for your websites.Test GenerationPython2026-02-27
61
PramodDutta
qaskills
140——14QA Skills Directory QA Skills is a curated directory of testing-specific skills for AI coding agents (Claude Code, Cursor, Copilot, etc.).E2E TestingTypeScript2026-06-08
62
alfredolopez80
multi-agent-ralph-loop
138——22Autonomous orchestration framework for Claude Code with MemPalace-inspired memory (4-layer stack, 818-token wake-up), parallel-first Agent Teams (6 teammates), Aristotle First Principles methodology, and 4-stage quality gates. 925+ tests, 22 active hooks, automatic learning pipeline.ObservabilityShell2026-06-07
63
Yenn503
Hexstrike-redteam
135——41 AI-powered MCP penetration testing framework combining HexStrike's 150+ security tools with BOAZ's advanced payload evasion (77+ loaders, 12 encoders). Features 12+ autonomous AI agents for bug bounty, CTF, CVE intelligence, exploit generation, and red team operations.Test GenerationC++2025-11-27
64
nbshenxm
pentest-agent
127——38PentestAgent is a novel LLM-driven penetration testing framework to automate intelligence gathering, vulnerability analysis, and exploitation stages, reducing manual intervention. For more information, read our paper at https://dl.acm.org/doi/10.1145/3708821.3733882 Security TestingPython2025-12-20
65
notadev-iamaura
OneRAG
124——38Production-ready RAG Framework (Python/FastAPI). 1-line config swaps: 6 Vector DBs (Weaviate, Pinecone, Qdrant, ChromaDB, pgvector, MongoDB), 5 LLMs (Gemini, OpenAI, Claude, Ollama, OpenRouter). OpenAI-compatible API. 2100+ tests.Test GenerationPython2026-06-10
66
Wytamma
write-the
121——15AI-powered Documentation and Test Generation ToolTest GenerationPython2025-06-26
67
Free-AI-Things
g4f-working
111——15g4f-working is a daily-updated list of working no-auth AI providers and models from @xtekky/gpt4free. It helps developers, testers, and AI enthusiasts instantly find which models are currently online and accessible without any API keys, tokens, or cookies.Test GenerationPython2026-06-09
68
qelos-io
testai
102——15The testing framework for skills, MCPs, commands, subagents, and LLM models! Docs at https://testingai.aiOtherTypeScript2026-05-25
69
Addepto
contextcheck
95——11 MIT-licensed Framework for LLMs, RAGs, Chatbots testing. Configurable via YAML and integrable into CI pipelines for automated testing.Test GenerationPython2024-12-11
70
vostride
agent-qa
93——4The self-improving Agentic QA harness with Memory. Write tests in natural language.
 Catch regressions before releases ship.E2E TestingTypeScript2026-06-03
71
nayangoel
AIPentester
91——15AI-powered penetration testing assistant for Claude Code. Autonomous security testing with browser automation, Burp Suite integration, and intelligent threat modeling.E2E TestingPython2026-04-28
72
georgeguimaraes
tribunal
91——3LLM evaluation framework for Elixir: evaluate and test LLM outputs, detect hallucinations, measure response qualityOtherElixir2026-06-10
73
maikotrindade
mobile-tester-agent
89——15AI-powered test automation tool which allows developers to run mobile automated testsOtherKotlin2026-05-28
74
anton-abyzov
specweave
87——9Spec-driven development framework for AI coding agents. 100+ skills for Claude Code, Cursor, Copilot, Codex, Antigravity & more. CLI tools, verified skill certification via verified-skill.com, and autonomous execution.OtherTypeScript2026-03-11
75
bitDive
java-producer
87——4Low-overhead Java agent for the Autonomous Verification Layer. Captures real runtime traces, SQL queries, and HTTP payloads in Spring Boot apps to enable trace-based testing.ObservabilityJava2026-04-29
76
JetBrains-Research
TestSpark
87——26TestSpark - a plugin for generating unit tests. TestSpark natively integrates different AI-based test generation tools and techniques in the IDE. Started by SERG TU Delft. Currently under implementation by JetBrains Research (Software Testing Research) for research purposes.Test GenerationKotlin2026-02-07
77
vuejs-ai
vue-tui
83——1The Vue framework for terminal UIs. SFC & JSX, Yoga flexbox, HMR, and testing out of the box.OtherTypeScript2026-06-09
78
bitDive
mcp-server
81——18BitDive Model Context Protocol (MCP) server. The Autonomous Quality Loop for AI agents. Provides real runtime context, before/after trace comparison, and integration testing workflows.ObservabilityPython2026-05-30
79
MLSZHU
LLMSafetyBenchmark
80——58A comprehensive framework for assessing the security capabilities of large language models (LLMs) through multi-dimensional testing.Security TestingPython2025-05-15
80
samihalawa
visual-ui-debug-agent-mcp
78——7VUDA is an autonomous debugging agent that empowers AI models to visually analyze, test, and debug webE2E TestingJavaScript2025-12-09
81
MasoudJTehrani
PCLA
77——12PCLA: A framework for testing autonomous agents in the CARLA simulatorOtherJupyter Notebook2026-04-13
82
DataKitchen
dataops-testgen
75——7DataOps Data Quality TestGen is part of DataKitchen's Open Source Data Observability. DataOps TestGen delivers simple, fast data quality test generation and execution by data profiling,  new dataset hygiene review, AI generation of data quality validation tests, ongoing testing of data refreshes, & continuous anomaly monitoringTest GenerationPython2026-06-03
83
ellydee
acceptance-bench
75——1A robust LLM evaluation framework measuring acceptance vs refusal across difficulty levels. Features multi-prompt variation testing, temperature sweeping, and LLM-as-judge evaluation.OtherPython2026-04-13
84
Yigtwxx
awesome-rag-production
74——21A curated list of battle-tested tools, frameworks, and best practices for building scalable, production-grade Retrieval-Augmented Generation (RAG) systems.Test GenerationPython2026-06-09
85
hoangsonww
Agentic-AI-Pipeline
70——20🦾 A production‑ready research outreach AI agent that plans, discovers, reasons, uses tools, auto‑builds cited briefings, and drafts tailored emails with tool‑chaining, memory, tests, and turnkey Docker, AWS, Ansible & Terraform deploys. Bonus: An Agentic RAG System & a Coding Pipeline with multistep planning, self-critique, and autonomous agents.OtherPython2026-06-05
86
syntropix-ai
synthora
68——2Synthora is a lightweight, extensible framework for LLM-driven agents and ALM research. It provides the essential components to build, test, and evaluate agents, enabling you to assemble an agent with a single configuration file. Our goal is to minimize effort while delivering robust functionality.OtherPython2025-10-03
87
rextanka
zero-cost-cline
64——3Tired of paying for frontier AI models? Break free with this guide to 100% local, private LLM code generation on Apple Silicon. Optimized for 24GB+ Macs, it uses Ollama + Cline to build a disciplined, free agentic workflow. Learn to write system rules, configure custom Modelfiles, and run strict two-layer test suites.Test Generation—2026-05-31
88
Rahulec08
appium-mcp
64——10AI-powered mobile automation with Model Context Protocol (MCP) integration. Seamlessly control Android & iOS devices through Appium with intelligent visual element detection and recovery. Built for AI agents like Claude to perform complex mobile testing workflows.Visual TestingTypeScript2025-07-30
89
srvsngh99
genai-testing-journey
63——2052-week journey from QA/SDET to GenAI Testing - learning in public with weekly mini-projects, code, and honest documentation of struggles and wins.OtherPython2026-05-12
90
existential-birds
beagle
62——8Claude Code plugin marketplace: 145 framework-aware code-review skills plus AI-writing detection, doc and test-plan generation, architectural analysis, and git workflows — for Python, Go, Rust, Elixir, React, Remix, iOS/Swift, and AI frameworks. Installable for Codex and other agents too.Test GenerationShell2026-06-10
91
zachblume
autospec
61——8Autospec is an open-source AI agent that takes a web app URL and autonomously QAs it, and saves its passing specs as E2E test codeE2E TestingTypeScript2026-05-15
92
walker0012025
API-TestPilot
61——13API-TestPilot,Ai 驱动的高效接口测试用例生成模型 | AI-Driven Efficient API Test Case Generation ModelTest GenerationPython2025-09-10
93
ServiceNow
DoomArena
61——7DoomArena is a Framework for Testing AI Agents Against Evolving Security ThreatsE2E TestingPython2025-09-12
94
zakirkun
saber-ai
61——10A production-grade, autonomous, AI-powered penetration testing and security automation platform.Security TestingPython2026-06-04
95
who0xac
Pinakastra
61——7AI-powered pentesting framework with automated recon and exploitation. Multi-source subdomain discovery, active vuln testing (XSS/SQLi/SSRF/IDOR), AI-driven payload generation, local inference, structured reporting. For pentesters and bug bounty hunters.Test GenerationGo2025-12-27
96
kousen
OpenAIClient
60——30Demonstrates how to use Spring to access OpenAI restful web services without using the Spring AI project. Tests call ChatGPT for text, DALL-E for image generation, and Whisper for audio transcriptions.Test GenerationJava2024-09-10
97
Echoxiawan
AITestCase
60——19A生成测试用例:基于页面和需求文档内容结合自动生成测试用例,解决单个需求文档生成测试用例质量较差的问题。Test Case Generation: Automatically generate test cases based on the current page and requirement document content, solving the problem of poor quality test cases generated from a single requirement document.E2E TestingPython2026-05-12
98
coinse
droidagent
59——8DroidAgent: Intent-Driven Mobile GUI Testing with Autonomous LLM AgentsOtherJupyter Notebook2024-03-12
99
Alqemist-labs
ruby_llm-tribunal
57——2LLM evaluation framework for Ruby, powered by RubyLLM. Tribunal provides tools for evaluating and testing LLM outputs, detecting hallucinations, measuring response quality, and ensuring safety. Perfect for RAG systems, chatbots, and any LLM-powered application.OtherRuby2026-04-09
100
presidio-oss
factif-ai
56——27AI-powered computer control for automated testing. Factifai uses vision models (Claude, GPT-4o, Gemini) to interact with applications naturally - clicking, typing, and verifying results just like a human would.E2E TestingTypeScript2025-10-01