Key Concepts
- On-device machine learning: Running machine learning models directly on devices like phones, rather than relying on cloud servers.
- Google AI Edge: A brand encompassing Google's on-device machine learning technologies, including TensorFlow Lite, MediaPipe, and Model Explorer.
- TensorFlow Lite: Google's framework for running TensorFlow models on mobile and embedded devices.
- MediaPipe: Google's framework for building cross-platform, customizable ML solutions for live and streaming media.
- Model Explorer: A tool for visualizing and analyzing machine learning models.
- Generative AI: AI models that can generate new content, such as text, images, or audio.
- LLM (Large Language Model): A type of AI model trained on a massive amount of text data, capable of understanding and generating human-like text.
- Retrieval-Augmented Generation (RAG): A technique that enhances LLMs by allowing them to retrieve information from external sources before generating a response.
- Embeddings: Numerical representations of data (e.g., text) that capture its meaning and relationships.
- Vector Database: A database optimized for storing and querying vector embeddings.
- Gemma: Google's family of open-source LLMs.
- MediaPipe LLM Inference API: An API that allows developers to run LLMs on-device using MediaPipe.
- Google AI Edge Torch generative API: An API that allows developers to create their own generative model architecture in PyTorch and convert that to run-on device.
- Quantization: A technique for reducing the size and computational cost of machine learning models by using lower-precision data types.
- Project Gameface: A Google project that enables people with mobility issues to interact with computers using facial gestures.
Sachin Kotwani's Background and Journey
- Sachin Kotwani is a product manager at Google, specializing in on-device machine learning.
- He has a diverse background, having lived in Spain, Africa, and the US, and speaks four languages: English, Spanish, Hindi, and Sindhi.
- Sachin's interest in technology began early, at age seven, with a Sinclair Spectrum computer. He learned programming by copying code from books and experimenting with modifications.
- He initially studied business management in college but later added a dual major in computer science.
- Before joining Google, Sachin worked in healthcare IT, software, and consulting.
- He joined Google as a strategy and operations manager, leveraging his MBA background.
- Sachin's interest in mobile development led him to build his own vegetarian recipe app, which gained half a million downloads.
- This experience helped him transition into product management at Google, where he worked on Google Cloud, Google Play, and Firebase.
- A colleague encouraged him to explore machine learning, leading him to take Andrew Ng's deep learning course.
- This led to his involvement in launching ML Kit, a popular SDK for on-device machine learning.
Google I/O and Google AI Edge
- Google I/O is Google's annual developer conference, showcasing the latest technology and innovations.
- Sachin Kotwani presented a talk on Google AI Edge at Google I/O 2024.
- Google AI Edge is a brand that brings together Google's on-device machine learning technologies, such as TensorFlow Lite, MediaPipe, and Model Explorer, under a single umbrella.
- The benefits of on-device machine learning include:
- Reduced reliance on cloud infrastructure, lowering costs and maintenance overhead.
- Low latency and offline connectivity for certain use cases.
- Enhanced privacy by processing data directly on the device.
- Sachin believes that generative AI models will increasingly move to devices as they become more efficient and devices become more capable.
Advantages of On-Device Machine Learning
- Privacy: Sensitive data can be processed on the device without being sent to the cloud.
- Latency: Real-time processing is possible without the delay of sending data to a server. Example: Background blur in Google Meet.
- Offline Functionality: Applications can function even without an internet connection.
- Cost Efficiency: Reduces the need for cloud resources, lowering operational costs.
New Features Launched at Google I/O 2024
- MediaPipe LLM Inference API: Enables running large language models (LLMs) on-device, supporting models ranging from 1B to 7B parameters, including Gemma 2B and 7B.
- Support for PyTorch: Allows developers to convert PyTorch models to TensorFlow Lite format and run them on-device.
- Google AI Edge Torch generative API: Enables developers to create custom generative model architectures in PyTorch and convert them to run-on device.
- Model Explorer: A tool for visualizing and analyzing machine learning models, helping developers understand their structure and performance.
Retrieval-Augmented Generation (RAG) with Smaller LLMs
- Smaller LLMs have limited capacity to store information but retain reasoning capabilities.
- RAG allows these models to access external knowledge sources, such as documents or databases, to answer questions more effectively.
- Example: A utility worker in Africa using an on-device LLM with RAG to troubleshoot equipment repairs by querying a manual stored on the device.
- RAG involves loading the manual into a vector database, creating embeddings for different chunks of text, and then retrieving the most relevant chunk based on the similarity between the question embedding and the content embeddings.
Ethical Considerations and Google's AI Principles
- Sachin acknowledges the potential for technology to be misused and emphasizes the importance of being deliberate and purposeful in its use.
- He highlights Google's AI principles, which prioritize privacy, ethics, and the use of AI for good.
- Google has teams dedicated to ensuring that products are developed and used responsibly.
Conclusion
Sachin Kotwani's journey highlights the importance of diverse experiences and continuous learning in the field of AI. The launch of Google AI Edge and its associated features at Google I/O 2024 represents a significant step forward in on-device machine learning, enabling developers to build more powerful, private, and accessible AI applications. The combination of smaller LLMs with RAG offers exciting possibilities for solving real-world problems, particularly in areas with limited internet connectivity. While ethical considerations remain paramount, the potential benefits of on-device machine learning are vast and continue to expand.
AI summaries can miss context or contain errors. Check important details against the original video.





