392 repos
1 / 4
| # | Repo ↓ | Stars ↓ | 1d | 7d | Forks | Description | Category | Language | Updated |
|---|---|---|---|---|---|---|---|---|---|
| 1 | vxcontrol pentagi | 17,608 | — | — | 2,408 | Fully autonomous AI Agents system capable of performing complex penetration testing tasks | Security Testing | Go | 2026-05-31 |
| 2 | raga-ai-hub RagaAI-Catalyst | 16,170 | — | — | 3,600 | Python SDK for Agent AI Observability, Monitoring and Evaluation Framework. Includes features like agent, llm and tools tracing, debugging multi-agentic system, self-hosted dashboard and advanced analytics with timeline and execution graph view | Performance Testing | Python | 2026-02-11 |
| 3 | cvat-ai cvat | 16,017 | — | — | 3,702 | Computer Vision Annotation Tool (CVAT) is a leading platform for building high-quality visual datasets for vision AI. It offers open-source, cloud, and enterprise products, as well as labeling services, for image, video, and 3D annotation with AI-assisted labeling, quality assurance, team collaboration, analytics, and developer APIs. | Visual Testing | Python | 2026-06-10 |
| 4 | alibaba MNN | 14,384 | — | — | 2,218 | MNN is a blazing fast, lightweight deep learning framework, battle-tested by business-critical use cases in Alibaba. Full multimodal LLM Android App:[MNN-LLM-Android](./apps/Android/MnnLlmChat/README.md). MNN TaoAvatar Android - Local 3D Avatar Intelligence: apps/Android/Mnn3dAvatar/README.md | Other | C++ | 2026-03-05 |
| 5 | web-infra-dev midscene | 13,653 | — | — | 1,036 | AI-powered, vision-driven UI automation for every platform. | E2E Testing | TypeScript | 2026-06-10 |
| 6 | GreyDGL PentestGPT | 13,627 | — | — | 2,360 | Automated Penetration Testing Agentic Framework Powered by Large Language Models | Security Testing | Python | 2026-06-07 |
| 7 | evidentlyai evidently | 7,590 | — | — | 861 | Evidently is an open-source ML and LLM observability framework. Evaluate, test, and monitor any AI-powered system or data pipeline. From tabular data to Gen AI. 100+ metrics. | Test Generation | Jupyter Notebook | 2026-05-02 |
| 8 | Giskard-AI giskard-oss | 5,425 | — | — | 468 | 🐢 Open-Source Evaluation & Testing library for LLM Agents | Security Testing | Python | 2026-06-10 |
| 9 | qodo-ai qodo-cover | 5,424 | — | — | 520 | Qodo-Cover: An AI-Powered Tool for Automated Test Generation and Code Coverage Enhancement! 💻🤖🧪🐞 | Test Generation | Python | 2026-04-05 |
| 10 | GH05TCREW pentestagent | 2,623 | — | — | 519 | PentestAgent is an AI agent framework for black-box security testing, supporting bug bounty, red-team, and penetration testing workflows. | Security Testing | Python | 2026-05-27 |
| 11 | Armur-Ai Pentest-Swarm-AI | 1,771 | — | — | 358 | Autonomous penetration testing using a swarm of AI agents. Orchestrates recon, classification, exploitation, and reporting specialists with ReAct reasoning — supports bug bounty, continuous monitoring, and CTF modes. Built with Go, Claude API, and 7+ native security tools. | Security Testing | Go | 2026-06-06 |
| 12 | kubeshop testkube | 1,605 | — | — | 163 | ☸️ The Open Testing Platform for AI-Driven Engineering Teams | Observability | Go | 2026-06-10 |
| 13 | zakirkun guardian-cli | 1,450 | — | — | 305 | Guardian is a production-ready AI-powered penetration testing automation CLI tool that leverages Google Gemini and LangChain to orchestrate intelligent, step-by-step penetration testing workflows while maintaining ethical hacking standards. | Test Generation | Python | 2026-05-29 |
| 14 | JoasASantos NeuroSploit | 1,122 | — | — | 273 | NeuroSploit is an advanced, AI-powered penetration testing framework designed to automate and augment various aspects of offensive security operations. Leveraging the capabilities of large language models (LLMs). | Security Testing | Python | 2026-03-29 |
| 15 | test-zeus-ai testzeus-hercules | 1,040 | — | — | 162 | Hercules is the world’s first open-source testing agent, enabling UI, API, Security, Accessibility, and Visual validations – all without code or maintenance. Automate testing effortlessly and let Hercules handle the heavy lifting! ⚡ | E2E Testing | Python | 2026-06-03 |
| 16 | SanMuzZzZz LuaN1aoAgent | 1,015 | — | — | 158 | LuaN1aoAgent is a cognitive-driven AI hacker. It is a fully autonomous AI penetration testing agent, using dual-graph reasoning. | Security Testing | Python | 2026-04-13 |
| 17 | bug0inc passmark | 950 | — | — | 165 | The open-source Playwright library for AI browser regression testing with intelligent caching, auto-healing, and multi-model verification. | E2E Testing | TypeScript | 2026-06-08 |
| 18 | MGdaasLab WHartTest | 934 | — | — | 125 | WHartTest 是一款AI驱动的测试自动化平台,实现从需求到可执行测试用例的自动化生成与管理,帮助测试团队提升效率与覆盖率。 (WHartTest is an AI-driven test automation platform that automates the generation and management of executable test cases from requirements, helping testing teams improve efficiency and coverage.) | E2E Testing | Python | 2026-06-04 |
| 19 | PurpleAILAB Decepticon | 932 | — | — | 175 | Autonomous Multi-Agent Based Red Team Testing Service / AI hacker | Test Generation | Python | 2026-03-26 |
| 20 | CyberSecurityUP NeuroSploit | 910 | — | — | 234 | NeuroSploit is an advanced, AI-powered penetration testing framework designed to automate and augment various aspects of offensive security operations. Leveraging the capabilities of large language models (LLMs). | Security Testing | Python | 2026-02-24 |
| 21 | alumnium-hq alumnium | 906 | — | — | 92 | End-to-end testing with AI | E2E Testing | TypeScript | 2026-06-09 |
| 22 | langwatch scenario | 895 | — | — | 65 | Agentic testing for agentic codebases | Other | Python | 2026-06-10 |
| 23 | Agent-Field SWE-AF | 844 | — | — | 133 | Autonomous software engineering fleet of AI agents for production-grade PRs on AgentField: plan, code, test, and ship. | Test Generation | Python | 2026-06-01 |
| 24 | zakirkun deep-eye | 831 | — | — | 177 | An advanced AI-driven vulnerability scanner and penetration testing tool that integrates multiple AI providers (OpenAI, Grok, OLLAMA, Claude) with comprehensive security testing modules for automated bug hunting, intelligent payload generation, and professional reporting. | Test Generation | Python | 2026-02-03 |
| 25 | ScrapeGraphAI scrapecraft | 656 | — | — | 99 | 🤖 AI-powered web scraping editor with visual workflow builder. Build, test & deploy web scrapers using natural language. Powered by ScrapeGraphAI & LangGraph. | Visual Testing | Python | 2025-12-26 |
| 26 | pikpikcu airecon | 643 | — | — | 100 | AIRecon is an autonomous cybersecurity agent that combines a self-hosted Large Language Model (Ollama) with a Kali Linux Docker sandbox and a Textual TUI. It is designed to automate security assessments, penetration testing, and bug bounty reconnaissance — without any API keys or cloud dependency. | Security Testing | Python | 2026-04-23 |
| 27 | CopilotKit aimock | 623 | — | — | 41 | Mock everything your AI app talks to — LLM APIs, MCP, A2A, AG-UI, vector DBs, search. One package, one port, zero dependencies. | Other | TypeScript | 2026-06-10 |
| 28 | trueleaf apiflow | 603 | — | — | 99 | A modern API workspace that works both online and offline — combining API documentation, testing, mock, and AI-powered automation in one lightweight tool | Other | TypeScript | 2026-06-10 |
| 29 | ServiceNow AgentLab | 586 | — | — | 116 | AgentLab: An open-source framework for developing, testing, and benchmarking web agents on diverse tasks, designed for scalability and reproducibility. | Performance Testing | Python | 2026-03-17 |
| 30 | PacificAI langtest | 560 | — | — | 49 | Deliver safe & effective language models | Performance Testing | Python | 2026-04-22 |
| 31 | Pacific-AI-Corp langtest | 551 | — | — | 50 | Deliver safe & effective language models | Performance Testing | Python | 2026-02-19 |
| 32 | modal-labs devlooper | 471 | — | — | 31 | A program synthesis agent that autonomously fixes its output by running tests! | Other | Python | 2024-09-19 |
| 33 | RevoltDevScript Revolt-Script | 461 | — | — | 1 | Revolt is the #1 Edgenuity automation tool featuring Auto Quiz, Auto Essay, Auto Advance, AI-powered writing with humanization, Auto Vocabulary, Auto Journal, and more. A semi-AFK script that handles quizzes, tests, exams, projects, and essays with ease. Compatible with desktop and mobile. Visit https://revolt.ly to get started. | Other | — | 2026-04-16 |
| 34 | project-codeguard rules | 410 | — | — | 56 | Project CodeGuard is an AI model-agnostic security framework and ruleset that embeds secure-by-default practices into AI coding workflows (generation and review). It ships core security rules, translators for popular coding agents, and validators to test rule compliance. | Test Generation | Python | 2026-01-29 |
| 35 | SHAdd0WTAka Zen-Ai-Pentest | 391 | — | — | 66 | 🛡⚔️AI-Powered Penetration Testing Framework with automated vulnerability scanning, multi-agent system, and compliance reporting🛡⚔️ | Security Testing | Python | 2026-06-08 |
| 36 | Ktovoz Skiritai | 383 | — | — | 16 | AI-powered test automation framework that explores test paths with AI and generates replayable scripts for 30x faster execution. | Test Generation | Python | 2026-05-30 |
| 37 | z3n70 Frida-Script-Runner | 363 | — | — | 75 | Web-based Frida framework and toolkit for Android & iOS penetration testing, mobile security, and dynamic analysis, featuring AI-assisted Frida script generation. | Test Generation | JavaScript | 2026-01-26 |
| 38 | faiscadev fakecloud | 316 | — | — | 19 | Free, open-source AWS emulator. LocalStack alternative: 40 services, 2,935 operations, true 100% Smithy conformance (99,678/99,678 variants pass). No account, no auth token, no paid tier. | Integration Testing | Rust | 2026-06-10 |
| 39 | CyberStrikeus CyberStrike | 316 | — | — | 58 | AI-powered offensive security agent with 7,300+ actionable security skills. Autonomous pentesting powered by MITRE ATT&CK (2,000+ Atomic tests), CIS Benchmarks (1,500+ controls), OWASP, NIST. Lazy-loading, zero context pollution. Your AI red team. | Security Testing | TypeScript | 2026-06-09 |
| 40 | aielte-research HackSynth | 310 | — | — | 51 | LLM Agent and Evaluation Framework for Autonomous Penetration Testing | Security Testing | Python | 2025-06-24 |
| 41 | ai-dashboad flutter-skill | 284 | — | — | 38 | AI-powered E2E testing for 10 platforms. 253 MCP tools. Zero config. Works with Claude, Cursor, Windsurf, Copilot. Test Flutter, React Native, iOS, Android, Web, Electron, Tauri, KMP, .NET MAUI — all from natural language. | E2E Testing | Dart | 2026-05-21 |
| 42 | pensarai apex | 283 | — | — | 49 | AI-powered offensive security testing using autonomous agents, directly in your terminal. | Security Testing | TypeScript | 2026-06-09 |
| 43 | jd-opensource JoySafeter | 282 | — | — | 53 | 🚀 JoySafeter: An enterprise AI Agent Platform—Not just chatting. building、running、testing, and tracing autonomous Agent Teams with visual orchestration... | Visual Testing | Python | 2026-05-29 |
| 44 | capture0x AdStrike | 257 | — | — | 52 | AI-powered modular Active Directory red-team framework for authorized penetration testing, AD enumeration, attack-path analysis, Kerberos/ADCS workflows, reporting, operator automation, and MCP server integration. | Security Testing | Python | 2026-06-02 |
| 45 | Coff0xc AutoRedTeam-Orchestrator | 247 | — | — | 50 | Enterprise AI Red Team Platform | 企业级AI红队平台 | 132 MCP Tools | Pure Python Engines | SDK+CLI+MCP | Auto-Download sqlmap/nuclei/ffuf | Production C2 | LLM Enhanced | Docker Sandbox | SARIF CI/CD | 1980 Tests | Security Testing | Python | 2026-05-18 |
| 46 | gsd-build gsd-browser | 238 | — | — | 15 | A fast, native browser automation CLI built from the ground up for AI agents, powered by Chrome DevTools Protocol. 63 commands covering navigation, interaction, screenshots, accessibility, network mocking, visual diffing, test generation, and more — all from a single binary. | E2E Testing | Rust | 2026-05-03 |
| 47 | praetorian-inc augustus | 234 | — | — | 30 | LLM security testing framework for detecting prompt injection, jailbreaks, and adversarial attacks — 190+ probes, 28 providers, single Go binary | Security Testing | Go | 2026-06-10 |
| 48 | Corpus-OS corpusos | 224 | — | — | 7 | Open-source protocol suite standardizing LLM, Vector, Graph, and Embedding infrastructure across LangChain, LlamaIndex, AutoGen, CrewAI, Semantic Kernel, and MCP. 3,330+ conformance tests. One protocol. Any framework. Any provider. | Other | Python | 2026-05-08 |
| 49 | LLAMATOR-Core llamator | 214 | — | — | 20 | Red Teaming python-framework for testing chatbots and GenAI systems. | Security Testing | Python | 2026-05-20 |
| 50 | inhouseseo superseo-skills | 199 | — | — | 27 | 11 Claude skills for SEO: page audits, linkbuilding, article writing, E-E-A-T audits, semantic gap analysis, link building. Methodology from Koray Tuğberk, Kyle Roof, and Lily Ray, plus a generation-time anti-AI-slop ruleset. Production-tested at InhouseSEO | Test Generation | — | 2026-05-12 |
| 51 | Hellsender01 LLMMap | 195 | — | — | 22 | Automated prompt injection testing framework for LLM-integrated applications with dual-LLM architecture. | Other | Python | 2026-03-14 |
| 52 | DanielSuo117 velocitai | 183 | — | — | 0 | Next-generation UI automation harness powered by Python & Playwright. Chat with AI Agents to seamlessly generate, architect, and execute enterprise-grade UI test code. | E2E Testing | HTML | 2026-05-24 |
| 53 | KHenryAegis VulnBot | 175 | — | — | 29 | The repository of VulnBot: Autonomous Penetration Testing for A Multi-Agent Collaborative Framework. | Other | Python | 2025-04-07 |
| 54 | CatchTheTornado open-agents-builder | 173 | — | — | 36 | AI Agents are missing the UI! We're here to change it. Build Business AI Agents for your company: business workflows, API's, bookings, e-commerce, social commerce, b2b, CPQ, intake forms, NPS tests, made-to-order use cases | Other | TypeScript | 2025-12-08 |
| 55 | AISquare-Studio AISquare-Studio-QA | 166 | — | — | 9 | Setting up QA testing agents using playwright and crewAI | E2E Testing | Python | 2026-04-17 |
| 56 | fugazi test-automation-skills-agents | 161 | — | — | 27 | A practical library of agents, instructions, and skills designed specifically for QA Automation Engineers, focusing on production-oriented solutions. | Other | Java | 2026-05-25 |
| 57 | ShirleyRex assertify.io | 157 | — | — | 121 | AI-powered assistant that turns your project context into end-to-end test strategies, detailed scenarios, and ready-to-run boilerplate across popular frameworks. | E2E Testing | TypeScript | 2026-03-15 |
| 58 | DebugBase glance | 151 | — | — | 23 | AI-powered browser automation MCP server for Claude Code. Navigate, click, screenshot, test — all from your terminal. | E2E Testing | TypeScript | 2026-04-16 |
| 59 | pymc-labs decision-lab | 150 | — | — | 12 | Run tested, autonomous agent workflows on your data for meaningful decision-making | Other | Python | 2026-06-03 |
| 60 | fzn0x watchtower | 141 | — | — | 24 | Watchtower is a simple AI-powered penetration testing automation CLI tool that leverages LLMs and LangGraph to orchestrate agentic workflows that you can use to test your websites locally. Generate useful pentest reports for your websites. | Test Generation | Python | 2026-02-27 |
| 61 | PramodDutta qaskills | 140 | — | — | 14 | QA Skills Directory QA Skills is a curated directory of testing-specific skills for AI coding agents (Claude Code, Cursor, Copilot, etc.). | E2E Testing | TypeScript | 2026-06-08 |
| 62 | alfredolopez80 multi-agent-ralph-loop | 138 | — | — | 22 | Autonomous orchestration framework for Claude Code with MemPalace-inspired memory (4-layer stack, 818-token wake-up), parallel-first Agent Teams (6 teammates), Aristotle First Principles methodology, and 4-stage quality gates. 925+ tests, 22 active hooks, automatic learning pipeline. | Observability | Shell | 2026-06-07 |
| 63 | Yenn503 Hexstrike-redteam | 135 | — | — | 41 | AI-powered MCP penetration testing framework combining HexStrike's 150+ security tools with BOAZ's advanced payload evasion (77+ loaders, 12 encoders). Features 12+ autonomous AI agents for bug bounty, CTF, CVE intelligence, exploit generation, and red team operations. | Test Generation | C++ | 2025-11-27 |
| 64 | nbshenxm pentest-agent | 127 | — | — | 38 | PentestAgent is a novel LLM-driven penetration testing framework to automate intelligence gathering, vulnerability analysis, and exploitation stages, reducing manual intervention. For more information, read our paper at https://dl.acm.org/doi/10.1145/3708821.3733882 | Security Testing | Python | 2025-12-20 |
| 65 | notadev-iamaura OneRAG | 124 | — | — | 38 | Production-ready RAG Framework (Python/FastAPI). 1-line config swaps: 6 Vector DBs (Weaviate, Pinecone, Qdrant, ChromaDB, pgvector, MongoDB), 5 LLMs (Gemini, OpenAI, Claude, Ollama, OpenRouter). OpenAI-compatible API. 2100+ tests. | Test Generation | Python | 2026-06-10 |
| 66 | Wytamma write-the | 121 | — | — | 15 | AI-powered Documentation and Test Generation Tool | Test Generation | Python | 2025-06-26 |
| 67 | Free-AI-Things g4f-working | 111 | — | — | 15 | g4f-working is a daily-updated list of working no-auth AI providers and models from @xtekky/gpt4free. It helps developers, testers, and AI enthusiasts instantly find which models are currently online and accessible without any API keys, tokens, or cookies. | Test Generation | Python | 2026-06-09 |
| 68 | qelos-io testai | 102 | — | — | 15 | The testing framework for skills, MCPs, commands, subagents, and LLM models! Docs at https://testingai.ai | Other | TypeScript | 2026-05-25 |
| 69 | Addepto contextcheck | 95 | — | — | 11 | MIT-licensed Framework for LLMs, RAGs, Chatbots testing. Configurable via YAML and integrable into CI pipelines for automated testing. | Test Generation | Python | 2024-12-11 |
| 70 | vostride agent-qa | 93 | — | — | 4 | The self-improving Agentic QA harness with Memory. Write tests in natural language. Catch regressions before releases ship. | E2E Testing | TypeScript | 2026-06-03 |
| 71 | nayangoel AIPentester | 91 | — | — | 15 | AI-powered penetration testing assistant for Claude Code. Autonomous security testing with browser automation, Burp Suite integration, and intelligent threat modeling. | E2E Testing | Python | 2026-04-28 |
| 72 | georgeguimaraes tribunal | 91 | — | — | 3 | LLM evaluation framework for Elixir: evaluate and test LLM outputs, detect hallucinations, measure response quality | Other | Elixir | 2026-06-10 |
| 73 | maikotrindade mobile-tester-agent | 89 | — | — | 15 | AI-powered test automation tool which allows developers to run mobile automated tests | Other | Kotlin | 2026-05-28 |
| 74 | anton-abyzov specweave | 87 | — | — | 9 | Spec-driven development framework for AI coding agents. 100+ skills for Claude Code, Cursor, Copilot, Codex, Antigravity & more. CLI tools, verified skill certification via verified-skill.com, and autonomous execution. | Other | TypeScript | 2026-03-11 |
| 75 | bitDive java-producer | 87 | — | — | 4 | Low-overhead Java agent for the Autonomous Verification Layer. Captures real runtime traces, SQL queries, and HTTP payloads in Spring Boot apps to enable trace-based testing. | Observability | Java | 2026-04-29 |
| 76 | JetBrains-Research TestSpark | 87 | — | — | 26 | TestSpark - a plugin for generating unit tests. TestSpark natively integrates different AI-based test generation tools and techniques in the IDE. Started by SERG TU Delft. Currently under implementation by JetBrains Research (Software Testing Research) for research purposes. | Test Generation | Kotlin | 2026-02-07 |
| 77 | vuejs-ai vue-tui | 83 | — | — | 1 | The Vue framework for terminal UIs. SFC & JSX, Yoga flexbox, HMR, and testing out of the box. | Other | TypeScript | 2026-06-09 |
| 78 | bitDive mcp-server | 81 | — | — | 18 | BitDive Model Context Protocol (MCP) server. The Autonomous Quality Loop for AI agents. Provides real runtime context, before/after trace comparison, and integration testing workflows. | Observability | Python | 2026-05-30 |
| 79 | MLSZHU LLMSafetyBenchmark | 80 | — | — | 58 | A comprehensive framework for assessing the security capabilities of large language models (LLMs) through multi-dimensional testing. | Security Testing | Python | 2025-05-15 |
| 80 | samihalawa visual-ui-debug-agent-mcp | 78 | — | — | 7 | VUDA is an autonomous debugging agent that empowers AI models to visually analyze, test, and debug web | E2E Testing | JavaScript | 2025-12-09 |
| 81 | MasoudJTehrani PCLA | 77 | — | — | 12 | PCLA: A framework for testing autonomous agents in the CARLA simulator | Other | Jupyter Notebook | 2026-04-13 |
| 82 | DataKitchen dataops-testgen | 75 | — | — | 7 | DataOps Data Quality TestGen is part of DataKitchen's Open Source Data Observability. DataOps TestGen delivers simple, fast data quality test generation and execution by data profiling, new dataset hygiene review, AI generation of data quality validation tests, ongoing testing of data refreshes, & continuous anomaly monitoring | Test Generation | Python | 2026-06-03 |
| 83 | ellydee acceptance-bench | 75 | — | — | 1 | A robust LLM evaluation framework measuring acceptance vs refusal across difficulty levels. Features multi-prompt variation testing, temperature sweeping, and LLM-as-judge evaluation. | Other | Python | 2026-04-13 |
| 84 | Yigtwxx awesome-rag-production | 74 | — | — | 21 | A curated list of battle-tested tools, frameworks, and best practices for building scalable, production-grade Retrieval-Augmented Generation (RAG) systems. | Test Generation | Python | 2026-06-09 |
| 85 | hoangsonww Agentic-AI-Pipeline | 70 | — | — | 20 | 🦾 A production‑ready research outreach AI agent that plans, discovers, reasons, uses tools, auto‑builds cited briefings, and drafts tailored emails with tool‑chaining, memory, tests, and turnkey Docker, AWS, Ansible & Terraform deploys. Bonus: An Agentic RAG System & a Coding Pipeline with multistep planning, self-critique, and autonomous agents. | Other | Python | 2026-06-05 |
| 86 | syntropix-ai synthora | 68 | — | — | 2 | Synthora is a lightweight, extensible framework for LLM-driven agents and ALM research. It provides the essential components to build, test, and evaluate agents, enabling you to assemble an agent with a single configuration file. Our goal is to minimize effort while delivering robust functionality. | Other | Python | 2025-10-03 |
| 87 | rextanka zero-cost-cline | 64 | — | — | 3 | Tired of paying for frontier AI models? Break free with this guide to 100% local, private LLM code generation on Apple Silicon. Optimized for 24GB+ Macs, it uses Ollama + Cline to build a disciplined, free agentic workflow. Learn to write system rules, configure custom Modelfiles, and run strict two-layer test suites. | Test Generation | — | 2026-05-31 |
| 88 | Rahulec08 appium-mcp | 64 | — | — | 10 | AI-powered mobile automation with Model Context Protocol (MCP) integration. Seamlessly control Android & iOS devices through Appium with intelligent visual element detection and recovery. Built for AI agents like Claude to perform complex mobile testing workflows. | Visual Testing | TypeScript | 2025-07-30 |
| 89 | srvsngh99 genai-testing-journey | 63 | — | — | 20 | 52-week journey from QA/SDET to GenAI Testing - learning in public with weekly mini-projects, code, and honest documentation of struggles and wins. | Other | Python | 2026-05-12 |
| 90 | existential-birds beagle | 62 | — | — | 8 | Claude Code plugin marketplace: 145 framework-aware code-review skills plus AI-writing detection, doc and test-plan generation, architectural analysis, and git workflows — for Python, Go, Rust, Elixir, React, Remix, iOS/Swift, and AI frameworks. Installable for Codex and other agents too. | Test Generation | Shell | 2026-06-10 |
| 91 | zachblume autospec | 61 | — | — | 8 | Autospec is an open-source AI agent that takes a web app URL and autonomously QAs it, and saves its passing specs as E2E test code | E2E Testing | TypeScript | 2026-05-15 |
| 92 | walker0012025 API-TestPilot | 61 | — | — | 13 | API-TestPilot,Ai 驱动的高效接口测试用例生成模型 | AI-Driven Efficient API Test Case Generation Model | Test Generation | Python | 2025-09-10 |
| 93 | ServiceNow DoomArena | 61 | — | — | 7 | DoomArena is a Framework for Testing AI Agents Against Evolving Security Threats | E2E Testing | Python | 2025-09-12 |
| 94 | zakirkun saber-ai | 61 | — | — | 10 | A production-grade, autonomous, AI-powered penetration testing and security automation platform. | Security Testing | Python | 2026-06-04 |
| 95 | who0xac Pinakastra | 61 | — | — | 7 | AI-powered pentesting framework with automated recon and exploitation. Multi-source subdomain discovery, active vuln testing (XSS/SQLi/SSRF/IDOR), AI-driven payload generation, local inference, structured reporting. For pentesters and bug bounty hunters. | Test Generation | Go | 2025-12-27 |
| 96 | kousen OpenAIClient | 60 | — | — | 30 | Demonstrates how to use Spring to access OpenAI restful web services without using the Spring AI project. Tests call ChatGPT for text, DALL-E for image generation, and Whisper for audio transcriptions. | Test Generation | Java | 2024-09-10 |
| 97 | Echoxiawan AITestCase | 60 | — | — | 19 | A生成测试用例:基于页面和需求文档内容结合自动生成测试用例,解决单个需求文档生成测试用例质量较差的问题。Test Case Generation: Automatically generate test cases based on the current page and requirement document content, solving the problem of poor quality test cases generated from a single requirement document. | E2E Testing | Python | 2026-05-12 |
| 98 | coinse droidagent | 59 | — | — | 8 | DroidAgent: Intent-Driven Mobile GUI Testing with Autonomous LLM Agents | Other | Jupyter Notebook | 2024-03-12 |
| 99 | Alqemist-labs ruby_llm-tribunal | 57 | — | — | 2 | LLM evaluation framework for Ruby, powered by RubyLLM. Tribunal provides tools for evaluating and testing LLM outputs, detecting hallucinations, measuring response quality, and ensuring safety. Perfect for RAG systems, chatbots, and any LLM-powered application. | Other | Ruby | 2026-04-09 |
| 100 | presidio-oss factif-ai | 56 | — | — | 27 | AI-powered computer control for automated testing. Factifai uses vision models (Claude, GPT-4o, Gemini) to interact with applications naturally - clicking, typing, and verifying results just like a human would. | E2E Testing | TypeScript | 2025-10-01 |