Agents
Choose the right Agent specialist
Match your source and expected result to Vision, Hearing, Voice, Language or Imagination.
The Agents product presents five specialist capabilities through a Python client. Before you copy an example, decide what you are giving the agent and what should come back. That simple distinction helps you choose a useful starting point.
Start with the source
For an image or a screenshot you want to understand, read the Vision example, agent.look. For an audio recording you want as text, read Hearing, agent.listen. These examples start from different kinds of source material.
Do not confuse listening with speaking. Hearing reads audio; Voice, shown as agent.speak, starts with the words you want spoken. Choosing the specialist starts with the direction of the task.
Define what should come back
Language, shown as agent.run, handles a written request. Imagination, shown as agent.imagine, is presented for image and idea generation. Read the example associated with the capability before assuming that the same argument works for every specialist.
Write a small expected outcome beside your request: extract the text on one receipt, transcribe one voice note or rewrite one paragraph. A result you can compare with the source is easier to assess than a broad instruction such as “handle everything”.
Separate the example from your execution
The animations on the Agents page illustrate the capabilities. They are not processing your files or executing a task in your account. Use them to understand the kind of input and output being described.
For an actual call, follow the Python setup shown on the page, use your own account key and inspect what the client returns. Check the returned value before connecting it to another step in your application.
Cyber Academy
What to remember
- Match the specialist to the input and output.
- Start with a small result you can verify.
- The page’s animated examples are illustrations, not live runs.