- reasoning
Agent-RL / ReCall
GitHub is where people build software. More than 150 million people use GitHub to discover, fork, and contribute to over 420 million projects.
11 Sept 2026 GitHubVisit → - voice
Deepgram Voice Agent
A library of function calling you can use with the Voice Agent API - deepgram-devs/voice-agent-function-calling-examples
10 Sept 2026 GitHubVisit → - code
OpenDevin
🙌 OpenHands: AI-Driven Development. Contribute to OpenHands/OpenHands development by creating an account on GitHub.
10 Sept 2026 GitHubVisit → - general
Plasma AI
Plasma AI builds infrastructure for intelligence at scale.
9 Sept 2026 WebVisit → - code
Code Review Agent Benchmark
Software engineering agents have shown significant promise in writing code. As AI agents permeate code writing, and generate huge volumes of code automatically -- the matter of code quality comes front and centre. As the automatically generated code gets integrated into huge code-bases -- the issue of code review and broadly quality assurance becomes important. In this paper, we take a fresh look at the problem and curate a code review dataset for AI agents to work with. Our dataset called c-CRAB (pronounced see-crab) can evaluate agents for code review tasks. Specifically given a pull-request (which could be coming from code generation agents or humans), if a code review agent produces a review, our evaluation framework can asses the reviewing capability of the code review agents. Our evaluation framework is used to evaluate the state of the art today -- the open-source PR-agent, as well as commercial code review agents from Devin, Claude Code, and Codex. Our c-CRAB dataset is systematically constructed from human reviews -- given a human review of a pull request instance we generate corresponding tests to evaluate the code review agent generated reviews. Such a benchmark construction gives us several insights. Firstly, the existing review agents taken together can solve only around 40% of the c-CRAB tasks, indicating the potential to close this gap by future research. Secondly, we observe that the agent reviews often consider different aspects from the human reviews -- indicating the potential for human-agent collaboration for code review that could be deployed in future software teams. Last but not the least, the agent generated tests from our data-set act as a held out test-suite and hence quality gate for agent generated reviews. What this will mean for future collaboration of code generation agents, test generation agents and code review agents -- remains to be investigated.
9 Sept 2026 WebVisit → - reasoning
Agent Test-Time Scaling
There is a popular assumption baked into many agentic AI systems: if an agent doesn't succeed on the first attempt, just let it try again. Give it more turns. Sample more trajectories. Add a reflection step. More compute at inference time should mean...
9 Sept 2026 WebVisit → - content
Blotato
Blotato is the #1 social media API and MCP server designed for AI agents, ChatGPT, and Claude. Schedule posts, manage comments and DMs, create content, and track analytics.
9 Sept 2026 WebVisit → - code
Runcell
Runcell is an AI agent for Jupyter notebooks that automates writing Python code, executing cells, debugging, and explaining data analysis results in real time.
9 Sept 2026 WebVisit → - tool-use
Sprocket
An open-source AI agent for hardware and software development that can autonomously purchase items online based on user commands, streamlining the procurement process.
8 Sept 2026 Hacker NewsVisit → - data
WebArena
Contribute to Agent-Tools/awesome-autonomous-web development by creating an account on GitHub.
17 Aug 2026 GitHubVisit → - general
OpenAgent
🔓 The open-source autonomous agent LLM initiative 🔓 - OpenAgentLLM/OpenAgent
17 Aug 2026 GitHubVisit → - general
ReactAgent
The open-source React.js Autonomous LLM Agent. Contribute to eylonmiz/react-agent development by creating an account on GitHub.
17 Aug 2026 GitHubVisit →