Back to News & insightsEngineering

AI recommendations: cold starts, feedback loops, and useful discovery

Build recommendations around reader intent, meaningful feedback, and a fair opportunity for new content instead of treating clicks as the whole objective.

Editorial guide · Updated September 28, 2026 · 7 min read
A mechanical selector hovers above stacked silver records beside a separate translucent disc.

A recommendation system does more than predict what someone might click. It decides which items receive attention, which authors find an audience, and which parts of a collection remain effectively invisible. Those decisions shape the data used to train the next version of the system. The product and the model gradually influence each other.

This makes recommendation engineering a problem of objectives, measurement, and controlled discovery. Better embeddings or a larger ranking model can help, but they cannot resolve an unclear definition of usefulness. Before building a personalized feed, decide what a successful recommendation should help a person accomplish and which forms of feedback can provide evidence of that success.

Separate finding candidates from ordering them

Google's recommendation-system overview describes a common division into candidate generation, scoring, and reranking. The first stage finds a manageable set of possibilities, the second estimates their relevance, and the final stage can apply additional constraints such as diversity or previously expressed preferences. This separation helps identify where a useful item disappeared.

Read source on developers.google.com

If a new article never enters the candidate set, improving the final ranking model cannot rescue it. If the candidates are relevant but the first screen repeats the same topic, the issue may be the ordering or diversity policy. Inspect the pipeline by stage instead of attributing every disappointing feed to one mysterious algorithm.

For a small catalog, start with a simple candidate generator that is easy to inspect. Topic tags, language, reading level, and explicit exclusions may provide a strong foundation. Sophisticated ranking becomes more useful after the basic collection metadata is accurate enough to support meaningful choices.

Define usefulness for a particular reading session

Imagine a hypothetical educational publication with beginner tutorials, research explainers, and advanced engineering guides. A reader who opens an introductory Python lesson may want the next prerequisite lesson, an example project, or a short explanation of an unfamiliar concept. Those are different intents, even when they involve similar words.

Ask what the recommendation slot is for. A continue learning slot should prioritize progression and prerequisites. A related research slot may deliberately broaden the topic. A catch up slot might favor material published since the reader's last visit. Combining all these intentions into one score can produce a feed that feels incoherent.

Write the objective in language an editor can review. Help the reader take the next useful step is a starting point; identify the behaviors that would support that interpretation. A saved tutorial, a completed exercise, or an explicit helpful rating may provide different evidence from an immediate click followed by a quick return.

Cold start is several different problems

A new reader has little behavioral history. A new article has little exposure history. An established reader exploring an unfamiliar subject creates another kind of uncertainty. Treating all three as the same missing-data problem often leads to a generic popularity list that serves none of them especially well.

For new readers, offer a small number of optional interests and an easy way to change them. Do not require an elaborate onboarding survey before showing useful material. The current page and an explicitly chosen goal can provide enough context to begin. Keep the nonpersonalized path useful for people who prefer not to create a profile.

For new articles, use editorial metadata and content representations to provide an initial opportunity for discovery. Absence of clicks is not evidence of poor quality when the item has barely been shown. Keep new-content exposure visible in internal reporting so a ranking system does not quietly freeze the catalog around old successes.

A click is conditional on being shown

Behavioral data is shaped by the interface. An item near the top of a list has a different opportunity to receive attention from an item below the fold. A striking thumbnail can attract a click even when the article disappoints. A quiet but useful reference may be saved and revisited without generating a dramatic immediate signal.

Record impressions with enough context to interpret interactions: the slot, position, candidate set version, and relevant presentation conditions. This does not require storing every detail of a person's browsing life. Collect the minimum event information needed for the evaluation and define retention deliberately.

When comparing models offline, acknowledge that historical logs reflect the previous system's choices. They do not reveal how people would have responded to every item that was never displayed. A model trained to imitate observed clicks can inherit the earlier system's blind spots and make them appear like objective preferences.

Exploration needs a purpose and a limit

Some discovery requires showing items whose usefulness is uncertain. That does not mean filling the feed with random content. Define an exploration policy around plausible relevance, editorial quality, and the reader's stated goals. A beginner should not receive an advanced systems paper merely because the system wants more data about it.

Reserve a bounded opportunity for less-exposed material and monitor what happens. Compare the resulting experience with a stable baseline. If readers repeatedly dismiss a category, respect that feedback rather than continuously reintroducing it under the name of exploration. Learning about preferences should not become an excuse to ignore them.

Keep the consequences proportionate. An optional reading suggestion is a reasonable place to experiment carefully. A recommendation that determines access to a scarce opportunity would require a different level of governance and evaluation. The technical machinery may look similar while the human stakes differ substantially.

Diversity should be defined in useful terms

Showing five articles with different titles does not guarantee a diverse list. They may all explain the same model release in slightly different language. Conversely, two items about the same topic can serve different needs, such as an introductory explanation and a practical debugging exercise.

Define diversity across dimensions that help the reader: topic, depth, format, perspective, or stage in a learning path. Avoid using variety as a decorative target detached from relevance. A coherent sequence can intentionally stay on one subject while varying the kind of work the learner performs.

Review real recommendation sets with editors. Ask whether the set duplicates effort, skips a prerequisite, or creates a useful choice. This qualitative review can expose problems hidden by item-level relevance scores. The unit of experience is often the whole set on the screen, not an isolated article considered independently.

Evaluate both immediate and delayed outcomes

An offline evaluation can check whether known relevant items are retrieved and ranked well. It is useful for catching regressions, but it should not be mistaken for a complete measure of reader benefit. The candidate collection, labels, and historical exposure process all influence what the result means.

In a controlled product test, measure signals that match the slot's purpose. For a learning path, examine whether readers reach the next lesson and report that it fits their level. For research discovery, look at saves, return visits, and explicit feedback. Explain the limitations of each measure instead of collapsing everything into one engagement number.

Track negative outcomes too: repeated dismissals, redundant suggestions, dead links, and recommendations that ignore language or accessibility preferences. A modest rise in clicks is not a success if the system becomes frustrating for a meaningful group of readers. Review outcomes across different kinds of sessions rather than only in aggregate.

Give people understandable controls

A reader should be able to hide an item, reduce a topic, or reset personalization without learning the internals of the ranking system. Make the effect of each control clear. Hide this article and stop recommending this topic are different requests and should not silently perform the same action.

Provide a useful explanation when possible, such as related to your current lesson or continues the retrieval series. Avoid explanations that invent a personal trait or imply certainty about motivation. The system observes limited behavior; it does not know the reader's identity, ambitions, or ability from a handful of clicks.

Check that feedback survives across the relevant surfaces. If a reader dismisses an article in the feed and immediately sees it in every related-content panel, the product feels careless. A consistent feedback policy can improve trust more than another small gain in a ranking metric.

Maintain a recommendation release record

Record the candidate rules, model revision, reranking policy, and evaluation results together. When a feed changes unexpectedly, the team should be able to identify whether the cause was a new model, altered metadata, or a policy change. Keep a fallback that delivers a useful editorial selection during failures.

Periodically inspect which parts of the catalog receive no exposure. The reason may be poor metadata, unsuitable language, duplication, or a ranking policy that never gives new work a chance. These are editorial and product questions as much as machine-learning questions.

A strong recommendation system helps people navigate a collection while remaining open to correction. Its quality comes from a clear purpose, honest interpretation of feedback, and a deliberate balance between relevance and discovery. The model is one component of that relationship, and it works best when the surrounding product makes the relationship understandable.

An original editorial guide. Provider capabilities and documentation can change. Follow the linked sources and test the exact model or service before relying on it.