Skip to main content
Golux Group

AI agents, multi-agent systems and AI employees

AI agents that do a real job, inside limits you set

We design and build single agents, multi-agent systems and AI employees: software that holds a defined role, uses your tools and data, hands work to other agents, and asks a person before anything risky leaves the building. Built to be operated and measured, not demoed once.

Start with one job a person does today: the steps, the tools they open, and where they make a judgment call. That is the best possible brief.

How a project works

How it starts
A written brief or a file you already have, then a call. Both free.
Who does the work
Nikola, plus a named collaborator where the work needs one: named before they start, never a bench.
What you get
Artifacts you own and can hand to any team, including one that is not us.
Payment
Split across agreed milestones. Nothing hourly by surprise.

Why most agent projects stall

The demo works. Then it meets a Tuesday.

An agent that impresses in a meeting and an agent you can leave running are different pieces of software. The gap is almost never the model. It is everything around it.

No job description
"Automate sales" is not a role. An agent needs the same thing a new hire needs: inputs, outputs, the tools it may use, what done looks like, and when to escalate.
One giant prompt doing everything
Research, planning, writing and checking in a single call is how quality collapses. Separate agents with separate responsibilities can be tested, replaced and trusted one at a time.
Actions without permissions
An agent that can send email, issue refunds or edit records needs scoped credentials, rate limits and an approval step. Without them, one bad input becomes one bad afternoon.
Nobody measures it
Without a test set of real cases, every prompt change is a guess. We build the evaluation before the agent, so "better" is a number and not an opinion.
Costs nobody forecast
Loops, retries and long contexts turn a cheap call into an expensive habit. Every system we ship has a budget per task and a ceiling it cannot cross.
No trail when something goes wrong
When a customer asks why the agent did something, you need the inputs, the tools it called and the reasoning it recorded. That log is designed in, not bolted on.

What we build

From one focused agent to a team of them

Each one is scoped as a job with an owner, tested against real cases, and shipped with the controls you need to run it yourself.

  • Single-purpose agents: triage, research, drafting, data entry, follow-ups, reporting
  • AI employees: an agent with a defined role, tools, working hours, limits and a human manager
  • Multi-agent systems: planner, specialist and reviewer agents that hand work to each other
  • AI agency systems: a service line run by agents end to end, with people approving the output
  • Tool use and integrations: your CRM, inbox, calendar, database, documents and internal APIs
  • Retrieval over your own knowledge, with sources attached to every answer
  • Human-in-the-loop approval for anything that spends money, contacts a customer or changes a record
  • Evaluation suites, cost ceilings, rate limits, audit logs and a kill switch
  • A dashboard where your team sees what each agent did, what it cost and what it escalated

How it runs

One job first, proven, then the next

  1. 01

    Pick the job

    We map one role a person does today: the inputs, the steps, the tools and the judgment calls. We also say plainly if the job should not be automated.

  2. 02

    Write the job description

    Responsibilities, permissions, what done looks like, and the exact moments it must hand over to a person.

  3. 03

    Build the test set

    Real past cases with known good outcomes. This is how we, and you, will know the agent is ready.

  4. 04

    Build and measure

    The agent, or the team of agents, built against the test set until it meets the bar you agreed to.

  5. 05

    Run it supervised

    In production with a person approving its output. Approval steps are removed only where the numbers say they can be.

  6. 06

    Hand over or operate

    Your team runs it with the dashboard and the runbook, or we keep operating and improving it on a monthly arrangement.

When an agent is the wrong answer

  • The process is not written down anywhere and nobody agrees on it. Automating confusion produces faster confusion. We would start by defining the process.
  • A plain script or a form would do it. If the steps never change, ordinary software is cheaper, faster and easier to trust, and we will tell you so.
  • You want to replace your team overnight. AI employees take over defined parts of jobs, under supervision, and earn more scope as they prove themselves.
  • There is no budget to run and maintain it. Agents need monitoring, evaluation and updates as models change, like any production system.

Why take this from us

We run agents in our own products, so we know where they break

Merlin Studio splits strategy from execution: one agent works through the audience, the offer and the goal, another turns that direction into briefs, drafts and checks, and a person approves the final version. Work Transcendence interviews people with an agent that asks back instead of guessing. PigeonAtlas lets agents send email only to people who asked for it. All of it built by the same team that has shipped over 50 products.

50+
products this team has built and shipped
2
agents in Merlin Studio, strategy and execution kept apart
1
person approving before anything is published
$300–400
per hour for project work, or a monthly arrangement
See Merlin Studio

Questions

AI agents, in detail

What is the difference between an AI agent, a multi-agent system and an AI employee?
An agent is software that uses a language model to decide which steps to take and which tools to call to finish a task. A multi-agent system splits work across several agents with separate roles, for example a planner, a specialist and a reviewer. An AI employee is an agent, or a small team of them, given a defined role in your company: a job description, tools, limits, a human manager and a way to measure its work.
Which models and frameworks do you use?
Whichever fits the job and your constraints. We work with the major model providers and keep the system model-agnostic where we can, so you can switch when prices or quality change. The framework matters less than the evaluation, the permissions and the logs around it.
Can the agent act on its own, or does a person always approve?
Both, decided per action. Reading, drafting and summarizing usually run on their own. Anything that spends money, contacts a customer or changes a record starts behind an approval step and comes out from behind it only when the measured results justify it.
Is our data safe?
Agents get the narrowest access that lets them do the job, through scoped credentials. We choose providers and settings that do not train on your data, keep sensitive fields out of prompts where possible, and log every tool call so access can be audited.
What about the EU AI Act and similar rules?
Some uses, such as hiring, credit and access to essential services, are treated as high-risk and carry obligations around documentation, oversight and transparency. We flag this in the first conversation and design the human oversight and records in from the start.
How much does an agent cost to build and to run?
Project work is $300–$400 an hour with a 10 hours minimum, and the number for your job comes after the call, in writing. Running costs are estimated per task before we build, and every system ships with a spending ceiling.
Can you add agents to software we already have?
Yes, and it is often the best place to start. Agents work through your existing APIs, database and tools. If those are not ready for it, we tell you what has to change first.

Send it over

Describe the job you want an agent to do

Who does it today, what they open, what they produce, and where they make a judgment call. Attach a process document, a sample of the work, or screenshots of the tools.

We reply with whether this is a good fit for an agent, what we would build first, and what we would measure.

Already written it down? Attach it.

PDF, Markdown, Word, text or images. Up to five files, 10 MB each. A spec you already have beats anything you could type into the box above.

The engineering notes

What we find inside AI-built apps.

The findings from the apps we audit, the fixes that worked, and what each one cost, written by the person who did the work. No roundups, no reposts.

A few emails a month. No spam, unsubscribe any time.

Golux Group

Already a client? Golux Club
is where your project lives: tickets, approvals, files, one record.

Open