383 repos
1 / 4
| # | Repo ↓ | Stars ↓ | 1d | 7d | Forks | Description | Category | Language | Updated |
|---|---|---|---|---|---|---|---|---|---|
| 1 | vxcontrol pentagi | 17,415 | — | — | 2,380 | Fully autonomous AI Agents system capable of performing complex penetration testing tasks | Security Testing | Go | 2026-05-31 |
| 2 | raga-ai-hub RagaAI-Catalyst | 16,169 | — | — | 3,607 | Python SDK for Agent AI Observability, Monitoring and Evaluation Framework. Includes features like agent, llm and tools tracing, debugging multi-agentic system, self-hosted dashboard and advanced analytics with timeline and execution graph view | Performance Testing | Python | 2026-02-11 |
| 3 | cvat-ai cvat | 15,967 | — | — | 3,695 | Computer Vision Annotation Tool (CVAT) is a leading platform for building high-quality visual datasets for vision AI. It offers open-source, cloud, and enterprise products, as well as labeling services, for image, video, and 3D annotation with AI-assisted labeling, quality assurance, team collaboration, analytics, and developer APIs. | Visual Testing | Python | 2026-06-03 |
| 4 | alibaba MNN | 14,384 | — | — | 2,218 | MNN is a blazing fast, lightweight deep learning framework, battle-tested by business-critical use cases in Alibaba. Full multimodal LLM Android App:[MNN-LLM-Android](./apps/Android/MnnLlmChat/README.md). MNN TaoAvatar Android - Local 3D Avatar Intelligence: apps/Android/Mnn3dAvatar/README.md | Other | C++ | 2026-03-05 |
| 5 | web-infra-dev midscene | 13,558 | — | — | 1,025 | AI-powered, vision-driven UI automation for every platform. | E2E Testing | TypeScript | 2026-06-03 |
| 6 | GreyDGL PentestGPT | 13,457 | — | — | 2,327 | Automated Penetration Testing Agentic Framework Powered by Large Language Models | Security Testing | Python | 2026-02-23 |
| 7 | evidentlyai evidently | 7,566 | — | — | 857 | Evidently is an open-source ML and LLM observability framework. Evaluate, test, and monitor any AI-powered system or data pipeline. From tabular data to Gen AI. 100+ metrics. | Test Generation | Jupyter Notebook | 2026-05-02 |
| 8 | Giskard-AI giskard-oss | 5,417 | — | — | 467 | 🐢 Open-Source Evaluation & Testing library for LLM Agents | Security Testing | Python | 2026-06-03 |
| 9 | qodo-ai qodo-cover | 5,412 | — | — | 518 | Qodo-Cover: An AI-Powered Tool for Automated Test Generation and Code Coverage Enhancement! 💻🤖🧪🐞 | Test Generation | Python | 2026-04-05 |
| 10 | GH05TCREW pentestagent | 2,569 | — | — | 509 | PentestAgent is an AI agent framework for black-box security testing, supporting bug bounty, red-team, and penetration testing workflows. | Security Testing | Python | 2026-05-27 |
| 11 | Armur-Ai Pentest-Swarm-AI | 1,645 | — | — | 341 | Autonomous penetration testing using a swarm of AI agents. Orchestrates recon, classification, exploitation, and reporting specialists with ReAct reasoning — supports bug bounty, continuous monitoring, and CTF modes. Built with Go, Claude API, and 7+ native security tools. | Security Testing | Go | 2026-05-25 |
| 12 | kubeshop testkube | 1,600 | — | — | 164 | ☸️ The Open Testing Platform for AI-Driven Engineering Teams | Observability | Go | 2026-06-03 |
| 13 | zakirkun guardian-cli | 1,445 | — | — | 304 | Guardian is a production-ready AI-powered penetration testing automation CLI tool that leverages Google Gemini and LangChain to orchestrate intelligent, step-by-step penetration testing workflows while maintaining ethical hacking standards. | Test Generation | Python | 2026-05-29 |
| 14 | JoasASantos NeuroSploit | 1,118 | — | — | 270 | NeuroSploit is an advanced, AI-powered penetration testing framework designed to automate and augment various aspects of offensive security operations. Leveraging the capabilities of large language models (LLMs). | Security Testing | Python | 2026-03-29 |
| 15 | test-zeus-ai testzeus-hercules | 1,034 | — | — | 160 | Hercules is the world’s first open-source testing agent, enabling UI, API, Security, Accessibility, and Visual validations – all without code or maintenance. Automate testing effortlessly and let Hercules handle the heavy lifting! ⚡ | E2E Testing | Python | 2026-05-26 |
| 16 | SanMuzZzZz LuaN1aoAgent | 995 | — | — | 158 | LuaN1aoAgent is a cognitive-driven AI hacker. It is a fully autonomous AI penetration testing agent, using dual-graph reasoning. | Security Testing | Python | 2026-04-13 |
| 17 | PurpleAILAB Decepticon | 932 | — | — | 175 | Autonomous Multi-Agent Based Red Team Testing Service / AI hacker | Test Generation | Python | 2026-03-26 |
| 18 | MGdaasLab WHartTest | 925 | — | — | 123 | WHartTest 是一款AI驱动的测试自动化平台,实现从需求到可执行测试用例的自动化生成与管理,帮助测试团队提升效率与覆盖率。 (WHartTest is an AI-driven test automation platform that automates the generation and management of executable test cases from requirements, helping testing teams improve efficiency and coverage.) | E2E Testing | Python | 2026-06-03 |
| 19 | CyberSecurityUP NeuroSploit | 910 | — | — | 234 | NeuroSploit is an advanced, AI-powered penetration testing framework designed to automate and augment various aspects of offensive security operations. Leveraging the capabilities of large language models (LLMs). | Security Testing | Python | 2026-02-24 |
| 20 | bug0inc passmark | 897 | — | — | 163 | The open-source Playwright library for AI browser regression testing with intelligent caching, auto-healing, and multi-model verification. | E2E Testing | TypeScript | 2026-05-29 |
| 21 | langwatch scenario | 893 | — | — | 65 | Agentic testing for agentic codebases | Other | Python | 2026-06-02 |
| 22 | alumnium-hq alumnium | 891 | — | — | 90 | End-to-end testing with AI | E2E Testing | TypeScript | 2026-06-03 |
| 23 | zakirkun deep-eye | 831 | — | — | 177 | An advanced AI-driven vulnerability scanner and penetration testing tool that integrates multiple AI providers (OpenAI, Grok, OLLAMA, Claude) with comprehensive security testing modules for automated bug hunting, intelligent payload generation, and professional reporting. | Test Generation | Python | 2026-02-03 |
| 24 | Agent-Field SWE-AF | 831 | — | — | 133 | Autonomous software engineering fleet of AI agents for production-grade PRs on AgentField: plan, code, test, and ship. | Test Generation | Python | 2026-06-01 |
| 25 | ScrapeGraphAI scrapecraft | 651 | — | — | 98 | 🤖 AI-powered web scraping editor with visual workflow builder. Build, test & deploy web scrapers using natural language. Powered by ScrapeGraphAI & LangGraph. | Visual Testing | Python | 2025-12-26 |
| 26 | pikpikcu airecon | 637 | — | — | 99 | AIRecon is an autonomous cybersecurity agent that combines a self-hosted Large Language Model (Ollama) with a Kali Linux Docker sandbox and a Textual TUI. It is designed to automate security assessments, penetration testing, and bug bounty reconnaissance — without any API keys or cloud dependency. | Security Testing | Python | 2026-04-23 |
| 27 | CopilotKit aimock | 615 | — | — | 40 | Mock everything your AI app talks to — LLM APIs, MCP, A2A, AG-UI, vector DBs, search. One package, one port, zero dependencies. | Other | TypeScript | 2026-06-03 |
| 28 | trueleaf apiflow | 604 | — | — | 98 | A modern API workspace that works both online and offline — combining API documentation, testing, mock, and AI-powered automation in one lightweight tool | Other | TypeScript | 2026-05-31 |
| 29 | ServiceNow AgentLab | 585 | — | — | 115 | AgentLab: An open-source framework for developing, testing, and benchmarking web agents on diverse tasks, designed for scalability and reproducibility. | Performance Testing | Python | 2026-03-17 |
| 30 | PacificAI langtest | 559 | — | — | 49 | Deliver safe & effective language models | Performance Testing | Python | 2026-04-22 |
| 31 | Pacific-AI-Corp langtest | 551 | — | — | 50 | Deliver safe & effective language models | Performance Testing | Python | 2026-02-19 |
| 32 | modal-labs devlooper | 469 | — | — | 31 | A program synthesis agent that autonomously fixes its output by running tests! | Other | Python | 2024-09-19 |
| 33 | Ktovoz Skiritai | 419 | — | — | 17 | AI-powered test automation framework that explores test paths with AI and generates replayable scripts for 30x faster execution. | Test Generation | Python | 2026-05-30 |
| 34 | project-codeguard rules | 409 | — | — | 56 | Project CodeGuard is an AI model-agnostic security framework and ruleset that embeds secure-by-default practices into AI coding workflows (generation and review). It ships core security rules, translators for popular coding agents, and validators to test rule compliance. | Test Generation | Python | 2026-01-29 |
| 35 | RevoltDevScript Revolt-Script | 388 | — | — | 1 | Revolt is the #1 Edgenuity automation tool featuring Auto Quiz, Auto Essay, Auto Advance, AI-powered writing with humanization, Auto Vocabulary, Auto Journal, and more. A semi-AFK script that handles quizzes, tests, exams, projects, and essays with ease. Compatible with desktop and mobile. Visit https://revolt.ly to get started. | Other | — | 2026-04-16 |
| 36 | SHAdd0WTAka Zen-Ai-Pentest | 385 | — | — | 65 | 🛡⚔️AI-Powered Penetration Testing Framework with automated vulnerability scanning, multi-agent system, and compliance reporting🛡⚔️ | Security Testing | Python | 2026-05-29 |
| 37 | z3n70 Frida-Script-Runner | 362 | — | — | 75 | Web-based Frida framework and toolkit for Android & iOS penetration testing, mobile security, and dynamic analysis, featuring AI-assisted Frida script generation. | Test Generation | JavaScript | 2026-01-26 |
| 38 | faiscadev fakecloud | 311 | — | — | 19 | Free, open-source AWS emulator. LocalStack alternative: 40 services, 2,935 operations, true 100% Smithy conformance (99,678/99,678 variants pass). No account, no auth token, no paid tier. | Integration Testing | Rust | 2026-06-02 |
| 39 | aielte-research HackSynth | 310 | — | — | 51 | LLM Agent and Evaluation Framework for Autonomous Penetration Testing | Security Testing | Python | 2025-06-24 |
| 40 | CyberStrikeus CyberStrike | 298 | — | — | 57 | AI-powered offensive security agent with 7,300+ actionable security skills. Autonomous pentesting powered by MITRE ATT&CK (2,000+ Atomic tests), CIS Benchmarks (1,500+ controls), OWASP, NIST. Lazy-loading, zero context pollution. Your AI red team. | Security Testing | TypeScript | 2026-06-02 |
| 41 | pensarai apex | 283 | — | — | 49 | AI-powered offensive security testing using autonomous agents, directly in your terminal. | Security Testing | TypeScript | 2026-06-03 |
| 42 | jd-opensource JoySafeter | 280 | — | — | 53 | 🚀 JoySafeter: An enterprise AI Agent Platform—Not just chatting. building、running、testing, and tracing autonomous Agent Teams with visual orchestration... | Visual Testing | Python | 2026-05-29 |
| 43 | ai-dashboad flutter-skill | 278 | — | — | 36 | AI-powered E2E testing for 10 platforms. 253 MCP tools. Zero config. Works with Claude, Cursor, Windsurf, Copilot. Test Flutter, React Native, iOS, Android, Web, Electron, Tauri, KMP, .NET MAUI — all from natural language. | E2E Testing | Dart | 2026-05-21 |
| 44 | Coff0xc AutoRedTeam-Orchestrator | 237 | — | — | 52 | Enterprise AI Red Team Platform | 企业级AI红队平台 | 132 MCP Tools | Pure Python Engines | SDK+CLI+MCP | Auto-Download sqlmap/nuclei/ffuf | Production C2 | LLM Enhanced | Docker Sandbox | SARIF CI/CD | 1980 Tests | Security Testing | Python | 2026-05-18 |
| 45 | gsd-build gsd-browser | 234 | — | — | 14 | A fast, native browser automation CLI built from the ground up for AI agents, powered by Chrome DevTools Protocol. 63 commands covering navigation, interaction, screenshots, accessibility, network mocking, visual diffing, test generation, and more — all from a single binary. | E2E Testing | Rust | 2026-05-03 |
| 46 | Corpus-OS corpusos | 224 | — | — | 7 | Open-source protocol suite standardizing LLM, Vector, Graph, and Embedding infrastructure across LangChain, LlamaIndex, AutoGen, CrewAI, Semantic Kernel, and MCP. 3,330+ conformance tests. One protocol. Any framework. Any provider. | Other | Python | 2026-05-08 |
| 47 | praetorian-inc augustus | 223 | — | — | 28 | LLM security testing framework for detecting prompt injection, jailbreaks, and adversarial attacks — 190+ probes, 28 providers, single Go binary | Security Testing | Go | 2026-06-03 |
| 48 | capture0x AdStrike | 216 | — | — | 43 | AI-powered modular Active Directory red-team framework for authorized penetration testing, AD enumeration, attack-path analysis, Kerberos/ADCS workflows, reporting, operator automation, and MCP server integration. | Security Testing | Python | 2026-06-02 |
| 49 | LLAMATOR-Core llamator | 211 | — | — | 20 | Red Teaming python-framework for testing chatbots and GenAI systems. | Security Testing | Python | 2026-05-20 |
| 50 | inhouseseo superseo-skills | 195 | — | — | 27 | 11 Claude skills for SEO: page audits, linkbuilding, article writing, E-E-A-T audits, semantic gap analysis, link building. Methodology from Koray Tuğberk, Kyle Roof, and Lily Ray, plus a generation-time anti-AI-slop ruleset. Production-tested at InhouseSEO | Test Generation | — | 2026-05-12 |
| 51 | Hellsender01 LLMMap | 195 | — | — | 22 | Automated prompt injection testing framework for LLM-integrated applications with dual-LLM architecture. | Other | Python | 2026-03-14 |
| 52 | ShirleyRex assertify.io | 176 | — | — | 121 | AI-powered assistant that turns your project context into end-to-end test strategies, detailed scenarios, and ready-to-run boilerplate across popular frameworks. | E2E Testing | TypeScript | 2026-03-15 |
| 53 | CatchTheTornado open-agents-builder | 172 | — | — | 36 | AI Agents are missing the UI! We're here to change it. Build Business AI Agents for your company: business workflows, API's, bookings, e-commerce, social commerce, b2b, CPQ, intake forms, NPS tests, made-to-order use cases | Other | TypeScript | 2025-12-08 |
| 54 | KHenryAegis VulnBot | 168 | — | — | 28 | The repository of VulnBot: Autonomous Penetration Testing for A Multi-Agent Collaborative Framework. | Other | Python | 2025-04-07 |
| 55 | AISquare-Studio AISquare-Studio-QA | 167 | — | — | 9 | Setting up QA testing agents using playwright and crewAI | E2E Testing | Python | 2026-04-17 |
| 56 | fugazi test-automation-skills-agents | 157 | — | — | 27 | A practical library of agents, instructions, and skills designed specifically for QA Automation Engineers, focusing on production-oriented solutions. | Other | Java | 2026-05-25 |
| 57 | DanielSuo117 velocitai | 156 | — | — | 0 | Next-generation UI automation harness powered by Python & Playwright. Chat with AI Agents to seamlessly generate, architect, and execute enterprise-grade UI test code. | E2E Testing | HTML | 2026-05-24 |
| 58 | DebugBase glance | 153 | — | — | 23 | AI-powered browser automation MCP server for Claude Code. Navigate, click, screenshot, test — all from your terminal. | E2E Testing | TypeScript | 2026-04-16 |
| 59 | pymc-labs decision-lab | 145 | — | — | 12 | Run tested, autonomous agent workflows on your data for meaningful decision-making | Other | Python | 2026-06-03 |
| 60 | fzn0x watchtower | 141 | — | — | 24 | Watchtower is a simple AI-powered penetration testing automation CLI tool that leverages LLMs and LangGraph to orchestrate agentic workflows that you can use to test your websites locally. Generate useful pentest reports for your websites. | Test Generation | Python | 2026-02-27 |
| 61 | alfredolopez80 multi-agent-ralph-loop | 139 | — | — | 22 | Autonomous orchestration framework for Claude Code with MemPalace-inspired memory (4-layer stack, 818-token wake-up), parallel-first Agent Teams (6 teammates), Aristotle First Principles methodology, and 4-stage quality gates. 925+ tests, 22 active hooks, automatic learning pipeline. | Observability | Shell | 2026-04-20 |
| 62 | PramodDutta qaskills | 133 | — | — | 11 | QA Skills Directory QA Skills is a curated directory of testing-specific skills for AI coding agents (Claude Code, Cursor, Copilot, etc.). | E2E Testing | TypeScript | 2026-05-25 |
| 63 | nbshenxm pentest-agent | 126 | — | — | 37 | PentestAgent is a novel LLM-driven penetration testing framework to automate intelligence gathering, vulnerability analysis, and exploitation stages, reducing manual intervention. For more information, read our paper at https://dl.acm.org/doi/10.1145/3708821.3733882 | Security Testing | Python | 2025-12-20 |
| 64 | notadev-iamaura OneRAG | 123 | — | — | 38 | Production-ready RAG Framework (Python/FastAPI). 1-line config swaps: 6 Vector DBs (Weaviate, Pinecone, Qdrant, ChromaDB, pgvector, MongoDB), 5 LLMs (Gemini, OpenAI, Claude, Ollama, OpenRouter). OpenAI-compatible API. 2100+ tests. | Test Generation | Python | 2026-06-01 |
| 65 | Wytamma write-the | 121 | — | — | 15 | AI-powered Documentation and Test Generation Tool | Test Generation | Python | 2025-06-26 |
| 66 | Free-AI-Things g4f-working | 108 | — | — | 14 | g4f-working is a daily-updated list of working no-auth AI providers and models from @xtekky/gpt4free. It helps developers, testers, and AI enthusiasts instantly find which models are currently online and accessible without any API keys, tokens, or cookies. | Test Generation | Python | 2026-06-02 |
| 67 | qelos-io testai | 101 | — | — | 14 | The testing framework for skills, MCPs, commands, subagents, and LLM models! Docs at https://testingai.ai | Other | TypeScript | 2026-05-25 |
| 68 | Addepto contextcheck | 95 | — | — | 11 | MIT-licensed Framework for LLMs, RAGs, Chatbots testing. Configurable via YAML and integrable into CI pipelines for automated testing. | Test Generation | Python | 2024-12-11 |
| 69 | georgeguimaraes tribunal | 90 | — | — | 3 | LLM evaluation framework for Elixir: evaluate and test LLM outputs, detect hallucinations, measure response quality | Other | Elixir | 2026-06-01 |
| 70 | nayangoel AIPentester | 90 | — | — | 15 | AI-powered penetration testing assistant for Claude Code. Autonomous security testing with browser automation, Burp Suite integration, and intelligent threat modeling. | E2E Testing | Python | 2026-04-28 |
| 71 | maikotrindade mobile-tester-agent | 90 | — | — | 15 | AI-powered test automation tool which allows developers to run mobile automated tests | Other | Kotlin | 2026-05-28 |
| 72 | JetBrains-Research TestSpark | 87 | — | — | 26 | TestSpark - a plugin for generating unit tests. TestSpark natively integrates different AI-based test generation tools and techniques in the IDE. Started by SERG TU Delft. Currently under implementation by JetBrains Research (Software Testing Research) for research purposes. | Test Generation | Kotlin | 2026-02-07 |
| 73 | anton-abyzov specweave | 87 | — | — | 9 | Spec-driven development framework for AI coding agents. 100+ skills for Claude Code, Cursor, Copilot, Codex, Antigravity & more. CLI tools, verified skill certification via verified-skill.com, and autonomous execution. | Other | TypeScript | 2026-03-11 |
| 74 | bitDive java-producer | 87 | — | — | 4 | Low-overhead Java agent for the Autonomous Verification Layer. Captures real runtime traces, SQL queries, and HTTP payloads in Spring Boot apps to enable trace-based testing. | Observability | Java | 2026-04-29 |
| 75 | bitDive mcp-server | 81 | — | — | 18 | BitDive Model Context Protocol (MCP) server. The Autonomous Quality Loop for AI agents. Provides real runtime context, before/after trace comparison, and integration testing workflows. | Observability | Python | 2026-05-30 |
| 76 | MLSZHU LLMSafetyBenchmark | 80 | — | — | 58 | A comprehensive framework for assessing the security capabilities of large language models (LLMs) through multi-dimensional testing. | Security Testing | Python | 2025-05-15 |
| 77 | vuejs-ai vue-tui | 80 | — | — | 1 | The Vue framework for terminal UIs. SFC & JSX, Yoga flexbox, HMR, and testing out of the box. | Other | TypeScript | 2026-06-03 |
| 78 | samihalawa visual-ui-debug-agent-mcp | 78 | — | — | 7 | VUDA is an autonomous debugging agent that empowers AI models to visually analyze, test, and debug web | E2E Testing | JavaScript | 2025-12-09 |
| 79 | MasoudJTehrani PCLA | 77 | — | — | 12 | PCLA: A framework for testing autonomous agents in the CARLA simulator | Other | Jupyter Notebook | 2026-04-13 |
| 80 | ellydee acceptance-bench | 76 | — | — | 1 | A robust LLM evaluation framework measuring acceptance vs refusal across difficulty levels. Features multi-prompt variation testing, temperature sweeping, and LLM-as-judge evaluation. | Other | Python | 2026-04-13 |
| 81 | vostride agent-qa | 76 | — | — | 4 | The self-improving Agentic QA harness with Memory. Write tests in natural language. Catch regressions before releases ship. | E2E Testing | TypeScript | 2026-05-25 |
| 82 | DataKitchen dataops-testgen | 75 | — | — | 6 | DataOps Data Quality TestGen is part of DataKitchen's Open Source Data Observability. DataOps TestGen delivers simple, fast data quality test generation and execution by data profiling, new dataset hygiene review, AI generation of data quality validation tests, ongoing testing of data refreshes, & continuous anomaly monitoring | Test Generation | Python | 2026-06-03 |
| 83 | hoangsonww Agentic-AI-Pipeline | 69 | — | — | 19 | 🦾 A production‑ready research outreach AI agent that plans, discovers, reasons, uses tools, auto‑builds cited briefings, and drafts tailored emails with tool‑chaining, memory, tests, and turnkey Docker, AWS, Ansible & Terraform deploys. Bonus: An Agentic RAG System & a Coding Pipeline with multistep planning, self-critique, and autonomous agents. | Other | Python | 2026-06-02 |
| 84 | syntropix-ai synthora | 68 | — | — | 2 | Synthora is a lightweight, extensible framework for LLM-driven agents and ALM research. It provides the essential components to build, test, and evaluate agents, enabling you to assemble an agent with a single configuration file. Our goal is to minimize effort while delivering robust functionality. | Other | Python | 2025-10-03 |
| 85 | Yenn503 Hexstrike-redteam | 65 | — | — | 12 | AI-powered MCP penetration testing framework combining HexStrike's 150+ security tools with BOAZ's advanced payload evasion (77+ loaders, 12 encoders). Features 12+ autonomous AI agents for bug bounty, CTF, CVE intelligence, exploit generation, and red team operations. | Test Generation | C++ | 2025-11-27 |
| 86 | Rahulec08 appium-mcp | 64 | — | — | 10 | AI-powered mobile automation with Model Context Protocol (MCP) integration. Seamlessly control Android & iOS devices through Appium with intelligent visual element detection and recovery. Built for AI agents like Claude to perform complex mobile testing workflows. | Visual Testing | TypeScript | 2025-07-30 |
| 87 | srvsngh99 genai-testing-journey | 62 | — | — | 20 | 52-week journey from QA/SDET to GenAI Testing - learning in public with weekly mini-projects, code, and honest documentation of struggles and wins. | Other | Python | 2026-05-12 |
| 88 | ServiceNow DoomArena | 61 | — | — | 7 | DoomArena is a Framework for Testing AI Agents Against Evolving Security Threats | E2E Testing | Python | 2025-09-12 |
| 89 | walker0012025 API-TestPilot | 61 | — | — | 13 | API-TestPilot,Ai 驱动的高效接口测试用例生成模型 | AI-Driven Efficient API Test Case Generation Model | Test Generation | Python | 2025-09-10 |
| 90 | who0xac Pinakastra | 61 | — | — | 7 | AI-powered pentesting framework with automated recon and exploitation. Multi-source subdomain discovery, active vuln testing (XSS/SQLi/SSRF/IDOR), AI-driven payload generation, local inference, structured reporting. For pentesters and bug bounty hunters. | Test Generation | Go | 2025-12-27 |
| 91 | existential-birds beagle | 61 | — | — | 8 | Claude Code plugin marketplace: 145 framework-aware code-review skills plus AI-writing detection, doc and test-plan generation, architectural analysis, and git workflows — for Python, Go, Rust, Elixir, React, Remix, iOS/Swift, and AI frameworks. Installable for Codex and other agents too. | Test Generation | Shell | 2026-05-31 |
| 92 | zachblume autospec | 61 | — | — | 8 | Autospec is an open-source AI agent that takes a web app URL and autonomously QAs it, and saves its passing specs as E2E test code | E2E Testing | TypeScript | 2026-05-15 |
| 93 | zakirkun saber-ai | 61 | — | — | 10 | A production-grade, autonomous, AI-powered penetration testing and security automation platform. | Security Testing | Python | 2026-05-19 |
| 94 | kousen OpenAIClient | 60 | — | — | 30 | Demonstrates how to use Spring to access OpenAI restful web services without using the Spring AI project. Tests call ChatGPT for text, DALL-E for image generation, and Whisper for audio transcriptions. | Test Generation | Java | 2024-09-10 |
| 95 | coinse droidagent | 59 | — | — | 8 | DroidAgent: Intent-Driven Mobile GUI Testing with Autonomous LLM Agents | Other | Jupyter Notebook | 2024-03-12 |
| 96 | Echoxiawan AITestCase | 59 | — | — | 19 | A生成测试用例:基于页面和需求文档内容结合自动生成测试用例,解决单个需求文档生成测试用例质量较差的问题。Test Case Generation: Automatically generate test cases based on the current page and requirement document content, solving the problem of poor quality test cases generated from a single requirement document. | E2E Testing | Python | 2026-05-12 |
| 97 | Alqemist-labs ruby_llm-tribunal | 56 | — | — | 2 | LLM evaluation framework for Ruby, powered by RubyLLM. Tribunal provides tools for evaluating and testing LLM outputs, detecting hallucinations, measuring response quality, and ensuring safety. Perfect for RAG systems, chatbots, and any LLM-powered application. | Other | Ruby | 2026-04-09 |
| 98 | presidio-oss factif-ai | 56 | — | — | 27 | AI-powered computer control for automated testing. Factifai uses vision models (Claude, GPT-4o, Gemini) to interact with applications naturally - clicking, typing, and verifying results just like a human would. | E2E Testing | TypeScript | 2025-10-01 |
| 99 | angelomorgado CARLA-GymDrive | 54 | — | — | 13 | Autonomous driving episode generation for the Carla simulator in a gym environment. This framework makes it easy to create driving scenarios to train/test the agent. | Test Generation | Python | 2024-11-01 |
| 100 | MatterAIOrg matter-ai | 53 | — | — | 7 | Matter AI is open-source AI Code Reviewer Agent 🤖 for Code Review, Summary Generation, Bug Detection, Security Vulnerabilities and Tests Generation | Test Generation | TypeScript | 2025-06-29 |