Local or cloud: where should your AI agent run?
Choose an agent's location by mapping three things separately: where the model processes requests, where tools act on files and services, and where the agent process and scheduler stay running. A local interface can use a remote model, and a cloud tool environment does not move a laptop's scheduler with it.
AI-assisted editorial: researched and drafted with AI using the primary sources linked below. Examples are illustrative; they are not measured product results. Editorial policy.
Key takeaways
- Model hosting, tool execution and the process that starts a routine are separate placement decisions.
- Choose local execution for a task's actual access needs, and remote execution for a defined operational requirement rather than a cloud label.
- Test the intended setup with the client closed and a dependency unavailable before relying on it for unattended work.
Map the model, tools and scheduler separately
The question usually starts with a practical need: should a bot work on my laptop or on a server? Before choosing, draw the complete path of one task. Include the process receiving the request, the model endpoint, the tools reading or changing data, and the process responsible for starting scheduled work. These components can live in different places.
Hermes documents terminal backends separately from model-provider configuration. Its local terminal backend executes commands on the host, while its SSH backend executes them on a remote server. That is a choice about command execution. It does not, by itself, tell you where the model processes a prompt.
For a proposed document assistant, write down a concrete map: agent process on laptop A, model endpoint B, files in folder C, output in destination D. Use actual configured endpoints and storage locations when implementing it. A diagram containing only a laptop icon and a cloud icon hides the decisions that affect the work.
Sources: Hermes Agent Configuration.
A local interface does not establish a local data path
Ollama's FAQ distinguishes running a model locally from using its cloud-hosted models, where prompts and responses are processed by the service. It also documents a local-only setting that disables Ollama cloud models and web search. This is a useful example of why the application's name or the location of its window cannot settle the data question.
Our recommendation is to inspect each boundary in the task. If a local agent sends a document excerpt to a hosted model, that excerpt crosses a network boundary. If a local model asks a connected search tool a question containing private project details, the tool request is a separate boundary to review. Turning off one application's cloud feature does not configure every other tool in an agent's workflow.
Decide which material the task needs before choosing its placement. A bot that sorts filenames may need much less information than one that summarizes complete documents. Use representative, non-sensitive fixtures to check the actual requests and outputs. Keep the data decision tied to the task rather than assuming that either local or cloud automatically meets it.
Sources: FAQ.
Place tools where their required resources are available
A tool needs a deliberate route to the files or service it uses. Consider a fictional bot that checks a folder on your workstation and creates a summary for your team. Moving its command execution to a remote machine does not make that folder appear there. You need a defined copy, mount or authorized connection, and you need to know where the result will be saved.
Hermes offers both container and remote terminal backends, among other options. Its configuration documentation describes the local backend as direct execution on the machine and the Docker backend as execution in a container. Treat that distinction as part of the access design, not a blanket promise about everything an agent can reach.
For our folder example, make a small resource inventory: one input directory, one output directory and any external API the routine actually needs. Verify the intended resources are accessible and unrelated resources are not. If the only way the design works is to give a simple sorting task broad access to an entire workstation, reconsider the task boundary before deciding which machine should host it.
Sources: Hermes Agent Configuration.
Put the scheduler where it can actually run
Hermes's scheduled-task documentation says its gateway daemon checks for due jobs and starts agent sessions. This makes the gateway process an operational dependency of Hermes scheduling. Choosing a remote terminal backend is a separate setting; it does not mean the gateway itself has moved to that remote environment.
In an illustrative hybrid setup, a laptop runs the gateway and sends commands to a remote machine. If the laptop stops running the gateway, the remote machine's availability alone is not sufficient to start a new scheduled agent session. To design for work while the laptop is unavailable, identify every process needed to trigger and complete that work and place those processes deliberately.
Keep client closure distinct from stopping the service. A chat window may be disposable while a managed background service continues running. Test the exact installation you plan to use: close the client, observe whether the service remains active, and check whether a scheduled fixture completes. Do not infer this behavior from a successful interactive conversation.
Sources: Scheduled Tasks (Cron).
Choose for one task before choosing for every bot
The following examples are proposed designs, not measured performance comparisons or claims about available BotBento features. Each begins with a different requirement. The point is to identify the dependency driving the decision, then test a small version before expanding it.
- Personal folder cleanup while you are present: start by evaluating a local tool environment with a narrow input folder and a reviewable output. Decide model hosting separately according to the content it will receive.
- A public-source briefing needed while your laptop is off: evaluate a remotely hosted agent process and scheduler, with only the source and delivery access required for that briefing.
- A team task that also needs a workstation-only file: evaluate a hybrid design, but make the workstation dependency visible. Define whether an unavailable file postpones the task or permits a clearly marked partial result.
- An occasional large computation: evaluate a separate execution environment for that step. Specify how inputs arrive and how outputs are retained before the temporary environment is removed.
Use a short placement acceptance checklist
For your chosen design, ask someone else to identify where a prompt is processed, where a command runs and where the schedule is owned. They should also be able to locate the output after the client is closed. If those answers require guessing, the setup is not yet clear enough to depend on.
Run a disposable fixture with one unavailable dependency. Inspect whether the outcome names that dependency and preserves any useful partial result. Then restore it and check the recovery behavior. Keep the evidence with the task's result record, including which parts were verified and which still need review.
Finally, name who maintains each running component: updates, credentials, storage and recovery need an owner in both local and remote arrangements. Compare measured resource use and real operational effort only after trying the representative task. This guide offers no universal price or speed winner. BotBento remains in development; these placement principles can help you evaluate an agent setup independently of its future interface.
Primary sources
Sources checked 2026-09-08. Standards and product documentation can change; follow the linked version when implementing.
- Hermes Agent Configuration — Nous Research
- FAQ — Ollama
- Scheduled Tasks (Cron) — Nous Research
BotBento is in development. Suggest a correction.