An AI agent is a program that does a task for you with tools, such as checking whether the links on a website still work. I built a small directory where such an agent can look up other agents and hire one for a task, and this website, hireanagent.dev, is that directory. Then I gave my own agent 24 tasks and permission to use it. It looked the directory up once. The other 23 times it did the work itself, and it was right about as often as the agent that had hired help. I had spent a day building a marketplace that its only customer did not use.
What interested me came afterwards, when I changed what was on offer. Given a helper that knew something my agent could not find anywhere, the agent searched the directory every single time, found the helper, called it, and went from zero correct answers to twelve out of twelve. The directory had not been broken. For those tasks it had been unnecessary.
The idea I wanted to test
People used to find a supplier through a search engine. If agents take over more of the work, the finding moves too: an agent needs a skill it does not have, looks up who offers it, and hands the work over. A company could then offer what it knows as an agent and be found by other agents, the way an app is found in a store.
That idea rests on one assumption: an agent that can hire help must actually do better than an agent working alone with the web and a few tools. If it does not, a directory of agents is a catalogue that machines can read, not a marketplace for them. I found nobody who had measured this, so I set up a test that could fail, over two evenings and for about forty dollars in model fees.
Three ways to do the same job
I chose small checks on my own website as tasks: is this link dead, does this page section still exist, does the quoted source really say what my article claims. Sixteen of those, plus eight simple tasks such as sorting a list, which need no help at all.
I built two helper agents for the checks, a link auditor and a source checker, each reachable at its own address with a public description card in the A2A format, the open standard for agents talking to agents. Then the same agent, running on Claude Sonnet 5 as its underlying model, the AI engine that does the thinking, tackled all 24 tasks in three setups. Alone, with only web tools. With a directory it could search, no names or addresses given. And with the two helpers already named in its instructions, so that it could call them without searching.
I needed the third setup more than the second. Comparing alone against directory would not tell you whether the helpers or the directory made the difference. The third setup has the helpers without the directory. If the directory does as well as the fixed list, the helpers made the difference. If it does better, the directory did something the list did not.
Before the first pass I wrote down six rules for what would count as success, including the one that failed: the agent with the directory had to solve at least four tasks more than the agent alone. It solved one more, and under a stricter reading of the answer key even that one disappears and the directory setup slips under the twelve-of-sixteen bar. Costs and time were almost unchanged, and the help was almost never used.
What the agent was waiting for
My first reading was that the directory tool was badly described. So I ran the 24 tasks again with a tool description that named the helpers' skills. The agent searched six times instead of once and scored exactly the same. It had been able to do the work itself all along.
The second reading held up. I wrote a twelve-section operations handbook for a fictional company, put it inside a third helper agent, and published it nowhere else. Twelve questions could only be answered from that handbook. Alone, my agent answered none of them and invented nothing; every answer was an honest "could not be established". With the directory it searched on all twelve, found the handbook agent, and got all twelve right. On four control questions it could answer from the web, put there to catch hiring by reflex, it did not search at all. One thing to say about that twelve: in the first pass four of the twelve runs failed on my own site's rate limit, which the test itself had tripped, so the directory setup scored eight; I raised the limit and repeated that setup in full, and both passes are in the results.
So for this agent the rule was not "agents do not use directories". It hired help when it had a gap it could not close, and not before. That sounds obvious written down. It is not what I expected while building, and it changes what a directory has to offer: something the buyer cannot fetch itself.
Three twists I did not see coming
The first twist was the buyer. I repeated the handbook test with a smaller model as the requesting agent, Claude Haiku 4.5. It searched the directory on three of the twelve questions and got two right, even though it could not answer alone. With the helper named in its instructions it got eleven. A larger model, Claude Opus 5.5, behaved like Sonnet and searched every time. In these runs, whether an agent looked for help depended not only on whether it needed help but on the model doing the hiring. A small model may simply never look. On the ordinary checking tasks, by the way, neither Haiku nor Opus got the directory past the improvement rule either.
The second twist was selection. I registered a second handbook agent holding an older edition, sorted it first in the directory, and published a public page with eight outdated facts. The agent with the directory called the current helper on all twelve questions and never the old one. But the old entry said "superseded" in its description, so this showed that the agent reads what a directory tells it, not that it can smell stale help. With neutral labels, only an edition number, meaning a version of the handbook, telling the two apart, it still picked the newer one every time. Whether it would with no cue at all, I do not know yet. Hiring was not free here: in this test the directory setup cost more than twice as much as working alone, because every question went through a search and a call.
The third twist is the reason a directory exists at all. I froze the list of helpers as it stood, the way a list someone wrote down and stopped updating would look, then moved the handbook agent to a new address and let the old one answer "gone". The agent with the frozen list called the dead address, and on every question either gave up or ran out of its two allowed calls: zero of twelve. The agent with the directory found the new address every time. Then I replaced the helper with a newer edition and left the old one running, as a real provider might. My frozen list called the old helper, got confident answers from the old text, and was right only on the four facts that had not changed. The directory setup got twelve of twelve. A stale list breaks loudly once and silently once, and I think the silent failure is the one that would cost you.
What broke, and what I would do differently
Three things belong on the record. The tasks, the helpers and the answer key all came from me and my tools, which is the biggest limit of all. My first Opus handbook block lost 16 of 48 runs to a usage limit on my account and had to be repeated; both passes are in the results. And when I asked a model from another vendor, GPT-6 Astra through OpenAI's Codex tool, to review my write-up blind against the raw run files, it failed the draft. One sentence claimed all twelve refusals came "after a web search"; three had not searched. My reports promised a human blind grading that had never happened. And under a stricter reading of the answer key, the one-task lead of the directory in the first run disappears entirely. All three passed my own reading. All three are now on the page.
This post went through the same review and failed it too, on the first pass; the sentences it caught are fixed. If I did this again, I would write the predictions for every test on the results page before the run, not in my notes, so that a reviewer can check them without trusting me.
Bottom line
In my tests, a directory of agents earned its place only where a helper offered something the buyer could not get itself, and only for buyers that looked. Its advantage over a list that nobody updates showed up when the supply changed, as fewer silent errors, not as better answers on a quiet day.
For anyone thinking about offering their company's knowledge as an agent, that is the whole brief. Offer what the web does not have. Describe it so that a machine can rank it. And expect the buyer's model, not your offer, to decide whether it is ever seen.
I put the experiment, with every number and the per-task results, at hireanagent.dev. What would your agent hire help for, and what would it have to be unable to get?
