API migration guide · Documentation review

Gemini API Model Migration: 3.7 Flash, 3.5 Flash and Deep Research Shutdown

Direct answer

Inventory your API model strings now, and migrate the old Deep Research agent before October 23, 2026. Google announced on October 8 that gemini-3.7-flash automatically routes to gemini-3.8-flash, and gemini-3.5-flash to gemini-3.6-flash. Those compatibility routes do not promise permanent support. The separate deep-research-pro-preview-12-2025 agent shuts down on October 23: explicitly change agent in interactions.create. Google API changelog.

AI Tool Finder Editorial Team · Sources checked October 11, 2026

This is a documentation-based guide, not a hands-on product test. Vendor capabilities are attributed below. Examples and verification steps are our editorial proposals; no account was connected and no product result is claimed.

The three migration decisions

Swipe wide tables horizontally on a phone.

Official API migration map — checked October 11, 2026
Old identifierCurrent replacementWhat you must decide
gemini-3.7-flashgemini-3.8-flashAutomatic routing is already announced. Name the target explicitly in your configuration and retest behavior.
gemini-3.5-flashgemini-3.6-flashAutomatic routing goes to 3.6, not 3.8. Moving directly to 3.8 is a separate model choice.
deep-research-pro-preview-12-2025deep-research-preview-04-2026 or deep-research-max-preview-04-2026Change the agent parameter. Select the speed-oriented or more comprehensive research workflow deliberately.

The changelog is about the Gemini developer API. A Gemini Apps subscription upgrade does not change your server configuration or prove API entitlement. Our consumer Gemini plan comparison owns that separate purchasing question.

October 8 announcement · Deprecation and shutdown definitions.

Deprecation, automatic routing and shutdown are different

October 8, 2026: Google deprecated the two Flash IDs and announced their automatic routes. As checked on October 11, we found no announced shutdown date for those two old Flash identifiers. Do not invent a deadline or interpret that absence as permanent support.

October 23, 2026: the old Deep Research preview is scheduled to shut down. Do not leave production creation requests on that identifier until the last day. The notice does not give a cutoff timezone in the cited entry; plan your migration before the date. We have not established how every already-running old-agent job will be handled at shutdown. Finish or review those jobs earlier and retain their IDs and outputs according to your data policy.

An old Flash request returning HTTP 200 can already be served by a different model. A successful transport response therefore proves neither unchanged behavior nor completion of your migration. Inventory the requested ID, any actual model-version metadata returned by your endpoint, SDK version and observed usage. Dated source.

Find the identifiers before changing them

rg -n --hidden -g '!node_modules/**' -g '!.git/**' -g '!vendor/**' 'gemini-3\.7-flash|gemini-3\.5-flash|deep-research-pro-preview-12-2025' .

Run the search in each application you own. Review matches in deployment configuration, environment-variable defaults, notebooks, queued-job producers, tests, prompt-routing tables and fallback lists. Search your deployment settings separately: repository search cannot see runtime-only configuration. Do not paste secret-bearing results into tickets.

  1. Classify each hit as a Flash model, a Deep Research agent, or a historical log/document. Preserve historical records.
  2. Record who owns the application, the endpoint and SDK version, current traffic, cost cap and rollback control. Stage the change in one configured route first.
  3. Replace the live configuration explicitly. Do not perform a blind global replacement: a 3.5-to-3.6 route and a deliberate 3.8 evaluation are different decisions.
  4. Repeat the search, inspect the diff and check deployed environment values. Keep an approved fixture set for the regression checks below.

Flash example and compatibility checks

This minimal Python example follows the current Google google-genai Interactions quickstart. It assumes you configure your own authorized API credentials through the SDK's supported environment setup. We checked its syntax locally; we did not install credentials or execute a model request.

from google import genai

# Requires your own authorized project and current google-genai SDK.
client = genai.Client()
result = client.interactions.create(
    model="gemini-3.8-flash",  # old 3.7 Flash now routes here
    input="Return a short deployment review checklist.",
)
print(result.output_text)

For an application that was using gemini-3.5-flash, use gemini-3.6-flash to name its announced route. You do not need to redesign an existing supported Flash endpoint merely to copy this Interactions example; consult the migration guide for your existing API contract. 3.8 Flash quickstart · 3.6 Flash guidance.

One concrete 3.8 compatibility check: the documented thinking levels are low, medium and high, with medium the default. minimal is unsupported and causes an error. Audit inherited settings instead of assuming every old parameter is compatible. Google lists a 1M-token context window and 64k maximum output; those limits do not guarantee your longest input or parser will work unchanged. Current API changes.

Deep Research: choose the agent and preserve the job lifecycle

Google describes deep-research-preview-04-2026 as the faster choice for interactive, user-facing research, including progress streaming. The deep-research-max-preview-04-2026 variant prioritizes comprehensive research and automated synthesis. Both remain preview agents; they are not two spelling variants of an identical service. We have not benchmarked their quality, time or cost. Deep Research guide.

Deep Research uses the Interactions API, not generate_content, and requires background=True. The following example uses Google's documented create/get/status/result fields. We added a local polling deadline so an operator can recover a known job instead of waiting indefinitely.

import time
from google import genai

client = genai.Client()
job = client.interactions.create(
    input="Compare the sources for our approved research question.",
    agent="deep-research-preview-04-2026",
    background=True,
)
print(f"Save this interaction ID: {job.id}")
# Editorial polling bound: reaching it does not cancel server work.
deadline = time.monotonic() + 30 * 60
while time.monotonic() < deadline:
    result = client.interactions.get(job.id)
    print(f"Status: {result.status}")
    if result.status == "completed":
        print(result.steps[-1].content[0].text)
        break
    if result.status == "failed":
        raise RuntimeError(f"Research failed: {result.error}")
    time.sleep(10)
else:
    raise TimeoutError(f"Check existing interaction {job.id}; do not resubmit blindly")

To evaluate the comprehensive variant, change only the agent value to deep-research-max-preview-04-2026 initially and run the same acceptance pack. Pin and review your SDK version; a much older SDK or a parser expecting a different response shape is not validated by changing the agent string alone.

For streaming, follow the current guide's stream=True flow and event schema. Confirm that the UI can show progress, reconnect without displaying duplicate chunks, recognize completion and retrieve the final result. A progress event is not a finished report. Preserve the interaction ID on creation so a dropped connection does not trigger an accidental duplicate research job. Check all statuses your installed SDK exposes, including interruption/cancellation if supported by your chosen workflow. Interactions lifecycle.

The snippet's deadline only stops local polling: it does not cancel a billable server job. A network exception also does not prove creation failed. Inspect the saved job before retrying creation. Add your application's approved retry/backoff and status handling around these minimal examples; do not treat them as a production worker implementation.

A migration is complete when the application still does its job

Swipe wide tables horizontally on a phone.

Before/after acceptance record — proposed checks, not tests we performed
AreaRecord before rolloutCheck on the replacement
IdentityRequested model/agent, SDK and endpointInspect actual version metadata where supplied. If absent, record it as unavailable; do not invent a resolved version.
Structured outputRepresentative schemas and parser resultsValid JSON, required fields, enums, nulls, numeric precision and truncated-output handling.
Tool callingAllowed tools, arguments and approval policyCorrect function and argument schema; errors and multi-step loops; preserve approval before external side effects.
Multimodal inputApproved image/audio/video fixturesSupported formats and limits; content interpretation and errors. Do not substitute text-only success.
Long context / billingShort and long fixtures; usage and cache assumptionsInput/output/thinking usage, truncation and cache behavior. Count failed attempts in cost.
Safety behaviorExpected allowed and disallowed examplesRefusals, blocked content, empty output and the UI explanation; do not weaken safeguards to force parity.
Errors / retriesTimeout, quota and transient-failure behaviorSeparate invalid parameters from retryable failures; bounded backoff and duplicate-job prevention.
Latency / unit costApplication SLO and accepted-output unitMeasure end-to-end latency and cost per accepted result, not just token price or a single fast response.
Deep ResearchCreate ID, get/status, final report and stream UISuccessful completion, failed job, interrupted stream, resumable polling and source links. A created job is not a delivered answer.

Use the same approved task set and record model identifiers, date, settings, outputs, usage, errors and a human acceptance decision. Set the tolerable failure and cost limits before switching traffic. These are future operator checks; this article reports no paid calls or successful product migration.

Recheck the bill, quotas and rollback assumptions

The current 3.8 guide lists an introductory $0.75 input / $3.75 output per million tokens through December 31, 2026, changing to $1.50 / $7.50 from January 1, 2027. This is a dated standard token-price reference, not a complete quote for every modality, cache, tool, research agent or account. Consult the specific pricing row and your project limits before estimating spend. Do not apply Flash token prices directly to Deep Research without checking its billing terms. 3.8 introductory terms · Current pricing · Rate limits.

Track the request's usage metadata and your actual billing view. More thinking or a longer agent loop can offset a lower unit rate. RPM, TPM and other constraints depend on the model and project tier; an advertised model launch does not increase your quota automatically.

Rollback warning: changing back to an old Flash name will still invoke its automatic route, so it cannot reliably restore the old model's behavior. Keep a separately verified supported alternative, feature switch, queue hold or human fallback. After October 23, the old Deep Research agent is not a viable rollback target. If a replacement fails acceptance, stop new research creation or use a tested supported path instead of resubmitting to a retired ID.

Roll out in a small monitored share, compare errors, latency, schema failures and cost, then expand only after your acceptance criteria pass. Keep watching the official changelog: this page was checked October 11 and cannot promise future retirement or access conditions.

Next step: make the inventory and choose the research agent

Start with the read-only identifier search, assign an owner to every live match, and schedule the Deep Research replacement before October 23. Then validate the smallest supported change on your own approved fixtures and budget.

Google Gemini profile · Consumer Gemini plan comparison · API speed-tier cost decisions · Agent tool workflow choices.