Choosing an AI model for your project

Chrome for DevelopersAbout 5 min readDec 26, 2025Watch original
THE SUMMARYAI-generated

Choosing the Right AI Model for Your Project

Key Concepts:

  • Foundational Models: Large, general-purpose AI models trained on massive datasets. (e.g., Gemini, GPT, Claude)
  • Small Language Models (SLMs): Smaller, more efficient models often fine-tuned for specific tasks. (e.g., Gemma, Feries, Quen Coder)
  • Expert/Task-Specific Models: Highly specialized models designed for a narrow range of functions. (e.g., object detection, OCR)
  • Inference: The process of using a trained model to make predictions or generate outputs.
  • Overfitting: When a model learns the training data too well and performs poorly on new data.
  • Underfitting: When a model is too simple to capture the underlying patterns in the data.
  • Data Gravity: The concept that data location influences where processing should occur.
  • Tokens: Units of text used in measuring the cost of using cloud-based language models.
  • Low-Rank Adaptation (LoRA): A technique for efficiently fine-tuning SLMs.
  • Web ML API: An API allowing web applications to interact with built-in language models in browsers.
  • Built-in AI: AI models integrated directly into browsers like Chrome and Edge.

1. Defining Needs and Constraints (Step One)

The initial step in selecting an AI model is a thorough understanding of project requirements. This involves clearly defining the tasks the model will perform, the problem it will solve, and identifying the target audience. Specificity is crucial. Key questions to answer include:

  • Decision/Action Support: What decision or action will the AI support?
  • Tasks Performed: What specific tasks will the model execute?
  • Expected Output: What form will the output take?
  • Output Usage: How will users interact with the output?
  • Success Metrics: What defines a successful outcome?
  • Target Audience: Who are the intended users?

Beyond functional requirements, technical goals and constraints must be considered. Guide tuning is recommended for testing model output. The speaker emphasizes that there is no universally “best” model; the optimal choice depends on data, problem-solving effectiveness, and deployment feasibility within technical and business limitations.

Constraints to consider:

  • Cost: Measured as cost per 10,000 tokens for cloud models, and bandwidth for client-side models. Cloud models may be more cost-effective for low prompt frequency, even with larger model sizes.
  • Data Security & Privacy: Regulations and requirements for local data access and storage.
  • Data Gravity: The location of the input data. Client-side models are advantageous if data resides on the client device.
  • System Integration: Compatibility with existing infrastructure.
  • Inference Performance: Real-time requirements versus acceptable latency (milliseconds or seconds). Asynchronous processing is an option for tasks not requiring immediate responses.
  • Device Limitations: Memory, disk space, and performance class of target devices for client-side deployment. Fallback mechanisms should be considered for devices not meeting minimum specifications.
  • Data Recency: How up-to-date the model’s training data needs to be.

2. Assessing Your Data (Step Two)

Before model selection, a comprehensive data assessment is vital. The quality, quantity, and structure of available data directly impact model performance and reliability. Insufficient or unsuitable data can lead to project failure. Understanding the data also helps determine model complexity and avoid overfitting or underfitting.

Assessment Factors:

  • Data Quality: Identifying missing information or inconsistencies.
  • Data Quantity: Determining if sufficient data exists for training and testing.
  • Data Type: Categorizing data as structured (e.g., spreadsheets), unstructured (e.g., text, images), or semi-structured (e.g., JSON). NLP is best suited for unstructured text.
  • Data Relevance: Assessing how well the data aligns with the intended use case.

If internal data is lacking, ethically sourced third-party data can be used, including open-source datasets, licensed databases, trusted vendors, or AI-generated synthetic data. Proper usage rights must be secured for external data.

3. Exploring Model Types (Step Three)

The speaker categorizes AI models into three main types:

  • Foundational Models: General-purpose models (e.g., Google Gemini, OpenAI GPT, Anthropic Claude) trained on massive datasets. They excel at complex reasoning and content creation but can be inefficient for specific tasks. Access is typically via APIs with subscription or per-token costs.
  • Small Language Models (SLMs): Smaller, more efficient models (e.g., Microsoft Feries, Google Gemma, Alibaba’s Quen Coder) fine-tuned for specific tasks. They offer a trade-off between flexibility and performance, with lower costs and resource requirements. Fine-tuning often utilizes Low-Rank Adaptation (LoRA). SLMs typically have 1 million to 10 billion parameters.
  • Expert/Task-Specific Models: Highly specialized models (e.g., object detection, OCR, text-to-speech) trained on domain-specific data for high accuracy and efficiency. These can be SLMs or other AI types.

Processing Location:

  • Server-Side (Remote/Cloud): Offers scalability, access to powerful hardware, and APIs.
  • Client-Side (Local): Provides data privacy, offline functionality, and reduces server load.
  • Hybrid: A combination of local and server-side processing.

The speaker notes that model type and processing location are often confused. Foundational models typically run on servers due to their size, but SLMs and expert models can be deployed on either servers or locally.

Real-World Example: Built-in AI in Chrome

Google is integrating SLMs and expert models directly into Chrome and other browsers. This "Built-in AI" allows websites to leverage AI functionality without self-hosting models. Websites interact with browser APIs to access local models (CPU, GPU, or NPU). Currently available APIs include translator, language detector, and summarizer, with more in origin trials. Microsoft Edge offers similar capabilities. The speaker highlights the standardization efforts for cross-browser compatibility. Chrome currently uses Gemini Nano and expert models for Built-in AI. Benefits include ease of deployment, hardware acceleration, and optimized performance.

Conclusion

Choosing the right AI model requires a systematic approach. The three-step process outlined – defining objectives and constraints, assessing data, and exploring model types – provides a framework for making informed decisions. The speaker encourages experimentation with different models, including built-in AI, to understand their capabilities and trade-offs. Resources like Chrome for Developers documentation offer further guidance and access to early access programs. Ultimately, the best model is the one that effectively solves the problem, fits the data, and aligns with technical and business constraints.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.