Laptop RAM 16GB đến 128GB thử chạy test gpt-oss-20B và -120B

Duy Luân Dễ ThươngAbout 5 min readAug 18, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

GPT-OSS (20B & 120B parameters), LM Studio, Token per second (token/s), Unify Memory (CPU & GPU share RAM), VRAM, Quantization, Cloud AI vs. Local AI, Privacy, RTX, Ryzen AI, Intel Arc Graphics.

GPT-OSS Model Testing on Various Laptops

Test Setup and Methodology

The video presents test results of running the GPT-OSS large language model (LLM) on various laptops, including both 20 billion (20B) and 120 billion (120B) parameter versions. The testing was conducted using LM Studio, updated to the latest version on both Windows and macOS. On Macs, the integrated graphics processing was used. On Windows, some tests used the dedicated GPU, while others used the CPU, depending on the situation. Windows laptops were tested while plugged in to ensure maximum performance. The same prompt was used across all devices. Results are compared against an RTX 4080 desktop for reference.

20B Parameter Model Performance

MacBooks, especially MacBook Pro models and Mac Studios, performed well, generally exceeding 30 tokens per second (token/s). A MacBook Pro M3 Pro (base model with 24GB RAM) was sufficient for running the 20B parameter model at a good speed. The advantage of Macs is their unified memory architecture, where RAM is shared between the CPU and GPU, and high RAM bandwidth. In the Windows laptop world, only AMD Ryzen AI Max Plus laptops (e.g., ROG Flow Z13 with 128GB RAM) have a similar unified memory architecture.

A token rate of 18-19 token/s is considered usable. Some Windows laptops without dedicated GPUs can achieve this. For example, an ASUS Vivobook TF414 (32GB RAM) achieved 14.49 token/s running live between GPU and CPU. Running on CPU alone, it reached 19.75 token/s. Running on the RTX 4060 alone resulted in only 10.52 token/s. This is because the 8GB VRAM of the RTX 4060 is insufficient to load the entire 20B parameter model, causing LM Studio to load some layers into system RAM, slowing down access.

MacBook Air M3 and M2 models with 16GB RAM can load the model, but performance is limited. LM Studio may require overriding a RAM limit. Even after loading, the system becomes slow and requires closing most other applications. 24GB RAM or more is recommended for a smooth experience. A MacBook Air M2 achieved 14.05 token/s, while an M3 achieved 16 token/s.

Intel Core Ultra 7 200V laptops (e.g., ThinkPad X1 Carbon Gen 12) are slow when processing with only the CPU. However, LM Studio now supports Intel's integrated graphics, increasing the token rate to 12.93 token/s (approximately 13 token/s), nearly three times faster than CPU-only processing. An ASUS Zenbook with an Intel Core Ultra 7 also saw similar improvements using integrated graphics (12 token/s vs 5 token/s on CPU).

High-end configurations like MacBook Pro M3 Max and ROG Flow Z13 with AMD Ryzen AI Max Plus can reach 40-70 token/s. The 20B model is suitable for code-related and SQL query tasks.

Nvidia GPUs are still the best for running LLMs, but most laptop RTX cards have insufficient VRAM. Only high-end cards like the RTX 4090 have enough VRAM (24GB) to load the model entirely on the GPU. Macs and high-end AMD laptops have the advantage of unified memory, allowing them to run the model, albeit slower, at a usable speed.

120B Parameter Model Performance

The MacBook Pro M3 Max performed the best. The RTX 4080 desktop could not load the model due to insufficient VRAM. Running the 120B model fully on a GPU requires 60-80GB of VRAM, which is only available on server or workstation cards. The AMD Ryzen AI Max Plus with 128GB RAM can run the model reasonably well, as can M3 Max and M3 Ultra Macs. When testing with the AMD Ryzen AI Max Plus, setting a fixed 96GB VRAM allocation caused LM Studio to report an error. Setting the VRAM allocation to "auto" resolved the issue.

Thermal Considerations

Running these models will cause the laptop fans to spin up. Macs have less fan noise. The GPU temperature on a MacBook Pro M3 Max can reach 100°C during continuous use, but it drops when idle. The speaker finds that using LM Studio for short bursts of questions is not problematic.

Cloud AI vs. Local AI: Is Local AI Necessary?

While running LLMs locally is possible, it is not necessary for most users. A subscription to a cloud-based AI service like ChatGPT Plus, Gemini, or Claude is sufficient for most needs. Cloud AI offers the advantage of processing power and can be accessed from any device, including smartphones. The cost (around $20/month) is reasonable. Buying a laptop specifically for running AI locally is not practical. Workstations are more suitable for dedicated AI tasks.

Cloud AI models are often larger, more accurate, and more intelligent than those that can be run locally. Cloud services are continuously updated and offer features like deep research and canvas mode (e.g., Gemini).

The trade-off with cloud AI is privacy. User data may be reviewed or used for model training, depending on the service's privacy policy. Running AI locally eliminates these privacy concerns.

The speaker recommends choosing the right tool for the situation. Cloud AI is sufficient for most users, but local AI is useful in situations where data cannot be sent to the cloud.

Laptop Buying Advice

When buying a laptop, the ability to run AI should not be a primary factor. Consider factors like battery life, aesthetics, screen quality, and keyboard/trackpad. AI performance should be a lower priority.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.