Profiles
Choose approved models, providers, local capacity, budgets, residency, and retention per Product or environment.
AI architecture
GlobalStacks separates AI application work from model-provider credentials and infrastructure details. Teams use a logical AI Gateway profile while policy can route requests to approved providers or private capacity on hosts they operate.
The operating model
A Product, environment, or AI agent selects one explicit gateway profile. That profile owns the approved models, provider connections, budget, residency and retention policy, and any private inference capacity that may serve the request.
Choose approved models, providers, local capacity, budgets, residency, and retention per Product or environment.
Keep provider credentials in the broker boundary. Workloads receive a gateway profile, never a provider key.
Resolve each request against one explicit profile, model alias, policy, and capacity decision.
Record content-minimized route evidence for operations, cost, and policy review without turning prompts into control-plane logs.
Request path
Local AI capability
No provider payment is required for requests served by your approved local capacity.
You still own the hardware, power, and operations cost. GlobalStacks makes that capacity governable rather than invisible.
Connected hosts can contribute approved CPU or accelerator capacity instead of forcing every request through a remote provider.
Inference workers stay behind typed agent and mesh paths. Their ports, management APIs, model files, and credentials are not application endpoints.
A profile can select an approved external provider, private local capacity, or a policy-defined fallback. The application keeps one gateway contract.
Clear boundaries
Choose a gateway profile and a model alias. They do not choose provider credentials, worker URLs, raw accelerators, or engine flags.
Applies model and provider policy, budgets, data handling, route selection, capacity admission, and route receipts.
Reports eligible capacity and runs approved private workers through typed lifecycle operations.
Move from the cloud to the edge
GlobalStacks does not ask a team to make one permanent infrastructure choice. A gateway profile gives the same application contract whether a request is served by approved provider capacity or a private worker on infrastructure you operate.
Keep inference near internal services, repositories, databases, and network boundaries instead of creating a new external data path for every request.
Put underused CPU or accelerator capacity to work with explicit scheduling and limits instead of paying a second runtime bill for steady workloads.
Avoid a round trip to a remote region when the application, retrieval source, or human workflow is already local to the selected host.
Use approved provider capacity for burst demand, models that do not fit your hardware, managed regional availability, or teams that do not want to operate inference workers. Profiles can make that choice explicit without changing application code or scattering provider keys.
The practical default
of recurring, well-bounded business requests can be a local-first design target.
Design for the work people actually repeat
Many business workflows are recurring and constrained: summarize a known report, classify a familiar request, retrieve from an approved knowledge set, prepare a standard response, or operate within a defined process. A well-designed prompt, tool boundary, and local model can serve that common path close to the data and at predictable cost.
The goal is not to force every imaginable question through a local model. Novel, ad hoc, or out-of-domain work can be routed to an approved larger model through the same gateway profile. That is an explicit exception path, with its own policy and receipt, not the architecture every routine request must inherit.
It is not a claim that a small local model can answer 80% of every question people could ask. It is a product-design hypothesis: identify the repeated requests in a specific business workflow, give them curated context and clear instructions, measure quality, and make that reliable path local-first. Keep escalation to larger models available for the work that does not fit.
Availability
Gateway profiles, provider connections, model aliases, route receipts, and private local execution foundations are in active delivery. Model-worker runtimes, accelerator lifecycle, scheduling, and additional engine support are enabled only where their typed host contract and verification coverage are complete.