Why AI with skills falls short at campaign optimization

AI with skills is very good at writing, coding and summarizing. In campaigns it struggles with data that stays incomplete for days, with algorithm learning phases and with its own unpredictability.

Phanes team9 min read
7 days of PRO free · no card

In text or code an AI mistake is visible right away and easy to fix. In a campaign a bad decision costs money, can restart the algorithm's learning and may only show up a week later. The stakes are different.

Picture a Monday. Someone exports the last seven days of results, pastes them into a chat with a campaign analysis skill and asks what to improve. The chat sees that a campaign with good ROAS "dropped" over the last two days and advises lowering its budget. But conversions from the last two days have not arrived yet, and the campaign has been in its learning phase for five days. The advice sounds reasonable and is wrong.

Where AI with skills works well
Why campaigns are different
The result is visible at once: text, code, a table.
The effect of a budget change shows after days or weeks.
A mistake is easy to undo and fix.
A change can restart the Google or Meta algorithm's learning phase.
The input data is complete.
Conversions from recent days are still arriving, so the data is incomplete.
A slightly different answer costs nothing.
A different answer on the same data means a different budget.
The whole problem fits in the context.
Account history is months of daily metrics across many campaigns.

Five traps in advertising data

Each is described in the official documentation of ad platforms and model providers. A chat analyzing an export sees none of them unless someone tells it.

Conversions by click date

Google attributes a conversion to the click date, and it can arrive up to 90 days later. The last few days always look worse than they are.

Data arrives with a delay

Clicks and cost appear within an hour, conversions after about 3 hours, and for some attribution models after about 15. Impression share refreshes once a day.

Learning phase

Google Smart Bidding usually learns for 1-2 weeks, sometimes up to 6. Meta needs about 50 events in the week after a significant edit. A hasty change restarts learning.

Limited context

The more tokens in the context, the lower the model's accuracy. Model vendors call this "context rot". A large account export is a textbook example.

Non-determinism

The same data can produce a different answer on the next call, even at temperature 0.

Any one of these traps can mislead an analysis. Together, they make a one-off look at an export likely to flag a problem that does not exist or to miss one that does.

How other tools make decisions

A chat with skills on an export

Someone downloads a report, pastes it into a chat with a skill and gets a conclusion. No history, no knowledge of learning phases, no measurement check.

Rules written by the team

A condition and an action, such as "if CPA exceeds X, pause". Predictable, but only as smart as the person who wrote them.

Statistical recommendations

Statistically significant patterns in account data, usually from one platform. Good for housekeeping, weaker at finding causes.

AI agent

A model analyzes the account and proposes changes itself. Flexible, but hard to trace why it did what it did.

Each approach has its place. They share one limitation: the decision depends on one person, one platform or one model call.

How Phanes decides

Phanes does not start by asking a model. It starts with data collected over time and with knowledge written into rules. The language model receives finished numbers and only explains them.

  1. Account history

    Every night Phanes stores each account's daily metrics, so the engine sees trends rather than a single export.

  2. Measurement verdict

    Google Ads, GTM, GA4 and the Meta pixel. Missing data is never treated as a clean result, and suspect measurement blocks conversion-based recommendations.

  3. 315 expert rules

    Rules from paid-ads expert practice and Google and Meta documentation, each with a condition, a threshold and a rationale.

  4. Learning-phase gate

    A campaign in learning gets no budget changes or restrictions. Learning window: 10 days on Google, 10 or 14 days on Meta.

  5. A hypothesis with a predicted effect

    Every proposal, and every change people make in the account, records a forecast: which metric should move, in which direction and over what period.

  6. Checked against reality

    After 7, 14 or 28 days, depending on the change type, the engine measures the result with volume and account-trend controls. Confirmations and failures weigh the same.

  7. Knowledge from many accounts

    Conclusions only from at least 3 accounts, none with more than a 40% share. Anonymized, without amounts or content.

Data, rules, approval: the path every change in the account follows.

Learning from data, not from prompts

Phanes learns the way a good specialist does: it writes down what it expects, then checks what happened. Every rule has a confidence score computed with a Bayesian method from settled trials. A rule that holds up gains priority. A rule that keeps missing loses it, up to and including being switched off. The engine also calibrates separately for each account and change type: where it hits, it acts more boldly; where it misses, it slows down.

Knowledge grows with every account while privacy stays intact. Cross-account conclusions come only from at least 3 accounts, none with more than a 40% share, and only relative indicators travel, never amounts or content. Before a recommendation reaches a card, the engine estimates a probability of improvement calibrated on the history of similar situations across many accounts. Data-driven knowledge never makes automation bolder without evidence from the account whose money is at stake.

Aether: a language model that only knows advertising

This is not an article against AI. Phanes has its own language model, Aether, built for one job: talking about campaigns using exactly what the engine computed. It does not know everything and does not need to. It is narrow, bilingual (Polish and English) and bound by Phanes's rules.

It describes engine verdicts

It turns the engine's findings into answers to four questions: what you see, since when, what it means and what to do. Every number comes from the input data.

It answers in "Ask Phanes"

It answers with evidence from the engine and a rule citation, checks the assumptions of your question and turns the request into an action that goes through the same preview and approval as any other change.

It classifies and structures

It recognizes off-offer, brand and generic search terms, extracts structure from a client brief or website, and translates texts between Polish and English.

The limits matter most. Aether makes no decisions about the account and never executes changes. It keeps no facts in its weights: it receives advertising knowledge from the Phanes knowledge base at question time, so updating knowledge needs no retraining. It learns form, discipline and style, and it does not learn to guess numbers.

  1. Data without names

    Training uses verdict descriptions, user decisions and expert labels after anonymization: no company names, people, addresses or account identifiers.

  2. Learning the format

    Light fine-tuning teaches the model the answer structure, data schemas and the dashboard language.

  3. Learning from verifiable rewards

    The model generates many answers to the same task, and automatic verifiers reward only those where every number comes from the data, the schema matches and the language is clean.

  4. Release gate

    A new version enters the product only if it beats the previous one on a frozen test set and does not get worse in any task category.

How Aether is built

Aether is not a single large model that knows everything. It is a small set of specialized components, each with one job and its own way of checking quality. The whole cycle, from training to answers in the product, runs on Phanes's own compute infrastructure: client data never leaves our servers and never goes to external model vendors.

What it isHow it works
Production modelA narrow, bilingual 7-14B-parameter model on open weights with a commercial license (Bielik and Qwen families)Describes verdicts, answers in the chat, classifies search terms, extracts structure from briefs, translates PL and EN. Served with plenty of throughput headroom, so it answers fast even across many accounts at once. The base is chosen by its score on the Phanes task set, not by public leaderboards.
Teacher and judgeA larger 27-32B-parameter model running locally, offline, in full precision (bf16)Generates training data variants from anonymized templates and grades answers. It never answers users.
Knowledge base (RAG)Multilingual embeddings (multilingual-e5) and an index of Phanes rules and knowledgeRetrieves the most relevant knowledge for every question. Facts live in the index, not in the weights, so new knowledge works after an index rebuild, without training.
Format tuning (SFT)LoRA, one epoch, low learning rate, verified data onlyTeaches the four-question answer structure, JSON schemas and the dashboard language.
Reinforcement learning (GRPO)Group Relative Policy Optimization with verifiable rewards, the method known from DeepSeek-R1For each task the model generates 8 answers. Verifiers score each one, and the model learns from the better ones in the group. Generation and training run in parallel on separate accelerators, and the 7-14B model trains in full precision (bf16) without quantizing the base. A KL penalty keeps it close to the base model so its style does not drift.
User preferences (DPO)Pairs: a card approved and a card rejected by the userAfter GRPO, the model learns how to describe recommendations that people actually accept. This step is evaluated separately.
Release gateA frozen evaluation set, separate for PL and ENThe target is 100% of answers without foreign numbers. A new version ships only if it beats the previous one with no regression in any category. Otherwise we roll back to the previous checkpoint.

The rewards Aether learns from

In reinforcement learning, everything depends on what the model is rewarded for. In Aether, every reward is checked mechanically, never by eye:

  • Numbers from the data: every number in the answer appears in the input data, in the same currency. Every foreign number is penalized.
  • Schema: the answer passes JSON structure validation.
  • Language: the whole answer is in the dashboard language, never mixing Polish and English.
  • Honesty: no promises of execution the engine does not provide, no knowledge source names, labels matching the expert, and a citation of a rule that exists.
  • Brevity: a penalty for verbosity. Training tasks are picked in the model's difficulty zone (20-80% correct answers from the base), because tasks that are too easy give no learning signal.

What the lab taught us

GRPO works and is safe

On GSM8K (1,319 problems) a small model gained 3.1 pp in 300 steps without changing its answer style, because the KL penalty kept it close to the base.

Training on its own samples can break a model

Fine-tuning on answers generated by the model itself cut the score by 15 pp. That is why every version passes a release gate, and rejection is a normal outcome.

A Polish base model wins in Polish

Bielik, a model trained on Polish data, beats larger English models on Polish tasks. Hence bilingual training and evaluation separately for PL and EN.

The biggest lesson is that the verifier matters more than the model. A single bug in the answer extractor understated scores by more than a percentage point and distorted comparisons, so every Phanes verifier has its own tests.

AI with skills is a good assistant. A decision about campaign money, though, should come from a system that remembers the account's history, knows about the learning phase, trusts only verified measurement and gives the same answer on the same data.

Why do the last few days in Google Ads look worse?

Google attributes a conversion to the click date, and the conversion itself can arrive up to 90 days later. On top of that, conversions show up in reports after about 3 hours, and for some attribution models after about 15. That is why the last few days always look worse than they are.

How long is the learning phase in Google Ads and Meta Ads?

Google Smart Bidding usually learns for 1 to 2 weeks, sometimes up to 6. Meta needs about 50 optimization events in the week after a significant edit. A hasty budget or settings change can restart learning.

What is Aether?

It is Phanes's own language model: a narrow, bilingual model trained only for Phanes tasks with GRPO and verifiable rewards. It describes the engine's conclusions and answers in the chat, but it makes no decisions and executes no changes. It runs on Phanes infrastructure and gets its knowledge from a knowledge base, not from its weights.

Sources (11)