I Tested GPT-5 as a Coding Agent—Here’s What Happened

Prompt EngineeringAbout 4 min readAug 18, 2025Watch original
THE SUMMARYAI-generated

Key Concepts:

  • GPT-5 (within Cursor): An agentic coding system.
  • Speech-to-text system: An application that transcribes speech into text.
  • MLX: Apple's Machine Learning eXchange framework for efficient model execution on Apple silicon.
  • Whisper models: Open-source speech recognition models by OpenAI.
  • Product Requirements Document (PRD): A document outlining the purpose, features, functionality, and behavior of a product.
  • Hotkeys: Keyboard shortcuts to trigger actions.
  • Virtual environment: An isolated environment for Python projects to manage dependencies.
  • LLM (Large Language Model): A deep learning model with a large number of parameters, trained to perform a variety of natural language processing tasks.
  • Quantization: A technique to reduce the memory footprint and increase the speed of neural networks by reducing the precision of the weights and activations.
  • Audible feedback: Audio cues to indicate the start and stop of recording.
  • Timeout limit: A set duration after which a process is automatically terminated.
  • Thinking tokens: Tokens generated by a language model during the reasoning process, which are not part of the final output.

1. Project Overview and Goals

  • The video demonstrates a real-world test of GPT-5 within the Cursor IDE by recreating a speech-to-text application.
  • The goal is to build a macOS-based application that transcribes speech to text, using MLX-optimized Whisper models.
  • The application should allow users to click in a text box, enable the app, and transcribe speech using hotkeys.
  • The presenter provides a detailed Product Requirements Document (PRD) and some code snippets to guide GPT-5.

2. Initial Implementation and Testing

  • GPT-5 generates a plan based on the PRD and creates a to-do list.
  • The initial implementation requires installing Python packages, which GPT-5 identifies.
  • The application is tested, and it successfully transcribes speech to text, displaying the transcribed text.
  • The initial test reveals minor issues, but the basic functionality works.
  • Example: "This was the quick recording of the word v."

3. Adding Model Selection Settings

  • The presenter requests the addition of settings to allow users to select from available Whisper models.
  • GPT-5 implements the feature, adding a settings panel with a list of models and an option to add custom models.
  • The custom model functionality initially appears to have issues, but it is later found to be a user error.
  • The presenter is able to add a custom model by providing a link to the model from the command line.

4. Implementing Audible Feedback

  • The presenter requests audible feedback when recording starts and stops.
  • GPT-5 implements the feature, adding sound cues for the start and stop of recording.
  • The implementation is tested and confirmed to be working.

5. Addressing Timeout Issues

  • The presenter identifies a timeout issue that stops transcription after a fixed duration.
  • GPT-5 is instructed to remove the timeout limit to allow for longer recordings.
  • The timeout limit is successfully removed.

6. Integrating an LLM for Error Correction

  • The presenter aims to integrate a small LLM to fix transcription errors.
  • GPT-5 uses its web search tool to identify suitable MLX-based LLMs, recommending "Qwen 1.7B".
  • The LLM is integrated as a secondary step to fix errors without rewriting the transcribed text.
  • The presenter instructs GPT-5 to ensure that the hotkey functionality remains the same.

7. Testing the LLM Integration

  • The integrated LLM is tested, and an error related to the "temperature" parameter is identified.
  • GPT-5 fixes the error.
  • The LLM integration is tested again, and it is found to be working, but it outputs "thinking tokens" during processing.
  • Example: "Here is a quick test. This is a quick test. I want the model to accurately identify the issues and fix those."

8. Conclusion and Future Work

  • The presenter concludes that GPT-5 successfully created a working speech-to-text application with error correction in a short amount of time.
  • The application has rough edges and bugs that need to be fixed.
  • The presenter plans to continue working on the application to create a more robust version.
  • The video demonstrates the potential of GPT-5 within Cursor for rapid application development.
  • Quote: "Overall the functionality works, transcription works as well as the fix of the transcribed text also works."
  • Quote: "It's pretty awesome that I was able to create a working app that people charge $20 per month within an hour."

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.