How Is This Legal? The Craziest AI Automation Yet (Wan 2.2 + n8n)

By Zubair Trabzada | AI Workshop

Share:

Key Concepts

  • WAN 2.2 Animate: A powerful AI model specifically designed for realistic face swap videos.
  • NADN: A no-code platform used for building and automating workflows, serving as the orchestrator for the face swap process.
  • NADN Blueprint: A pre-configured workflow file that can be imported into NADN to quickly set up complex automations.
  • Form Submission (NADN): A trigger node in NADN that allows users to upload files (like videos and images) and initiate a workflow.
  • Cloudinary: A digital asset management (DAM) service used to host and manage media files (images, videos) in the cloud, providing accessible URLs for API interactions.
  • HTTP Request Node: A NADN node used to send API requests to external services (e.g., Cloudinary, File.ai) for data exchange.
  • Set Node: A NADN node used to transform, clean, or extract specific data points from previous nodes' outputs, making them more usable for subsequent steps.
  • Merge Node: A NADN node that combines data from multiple input branches into a single output, streamlining data flow.
  • File.ai: A platform that provides API access to various advanced AI models, including WAN 2.2 Animate, simplifying integration.
  • API Endpoint: A specific URL that an API uses to receive requests and send responses.
  • API Key: A unique identifier used for authentication when accessing an API.
  • Resolution: The quality of the video output (e.g., 480p, 720p), which can impact processing cost and time.
  • Wait Node: A NADN node that pauses the workflow for a specified duration, often used when waiting for asynchronous processes (like video generation) to complete.
  • If Node: A NADN node that introduces conditional logic, allowing the workflow to branch based on specific criteria (e.g., checking if a job status is "completed").
  • Request ID: A unique identifier assigned to an API job, used to track its status and retrieve results.
  • Lip Sync: The synchronization of lip movements with spoken audio, a key indicator of realism in face swap technology.

Introduction to AI Face Swapping and Content Creation

The video introduces WAN 2.2 Animate as one of the most powerful AI models for creating "mind-blowing face swap videos completely on autopilot." It highlights the ease of use with NADN, a no-code platform, where users simply upload a video and an image for face swapping, and the AI handles the rest, generating "incredibly realistic results." The speaker emphasizes that this technology marks "a brand new era of content creation," making cinematic, high-quality videos "faster, easier, and more accessible than ever before." This shift levels the playing field for creators, business owners, and beginners, removing the need for coding, teams, or professional hires. The core message is: "You just need to start because now there is no excuse not to."

NADN Workflow Setup and Initial Input

The process begins by accessing resources from the "AI workshop community," specifically downloading the NADN blueprint. This blueprint is then imported into a free NADN account by clicking the three dots in the top right corner and selecting "import file." The workflow is executed by clicking "execute workflow," which brings up a form. This form allows the user to upload a source video and an image for the face swap. The speaker demonstrates this by uploading a 14-second sample video of himself talking about AI and an image of a viral Spanish AI model (not a real person, but an AI-generated image that gained popularity for generating significant income).

The Form Submission node in NADN acts as the trigger. It's configured with a description ("upload an image to swap the face in the video") and two "file" type form elements: one for the video and one for the image. This setup ensures that the uploaded files are correctly received by the workflow.

Cloudinary Integration for Asset Hosting

The next crucial step involves uploading the video and image to a cloud storage service to obtain publicly accessible URLs, as the WAN 2.2 Animate API requires URLs for its inputs. The video recommends Cloudinary, a digital asset management service, due to its "generous free account" which offers up to "2500 assets or 25 GB" of storage.

Two HTTP Request nodes are used to upload the video and image to Cloudinary. These nodes interact with Cloudinary's API endpoint (api.cloudinary.com/v1_1/auto/upload). The requests are sent as "form data" and include an "upload preset" name, which is configured within the user's Cloudinary account. After successful upload, Cloudinary provides a "secure URL" for each asset. The speaker demonstrates that the uploaded image and sample video appear in his Cloudinary assets, confirming their availability via URL.

Data Cleaning and Merging

Following the Cloudinary upload, two Set nodes are used to clean up the output. The raw output from the HTTP request nodes contains a lot of unnecessary data. The Set nodes extract only the "secure URL" for both the video and the image, renaming them simply "video" and "image" respectively. This makes the data cleaner and easier to use in subsequent steps.

A Merge node then combines these two cleaned outputs (the video URL and the image URL) into a single, structured output. This ensures that both necessary URLs are available together for the next stage of the workflow.

WAN 2.2 Animate API Interaction via File.ai

The core face swap operation is performed by sending an API request to WAN 2.2 Animate through File.ai. File.ai is presented as a platform that provides API access to various AI models. The specific endpoint used is https://api.file.ai/v1/models/wan-v2.2-14b-animate-replace/predict.

An HTTP Request node is configured to send a POST request to this endpoint. Key components of this request include:

  • Header Authorization: Requires an API key for authentication.
  • Body: Contains the essential parameters for the face swap:
    • video_url: The cleaned URL of the source video from Cloudinary.
    • image_url: The cleaned URL of the face swap image from Cloudinary.
    • resolution: Specifies the desired output resolution. The video mentions options like 480p, 580p, and 720p. The speaker chooses 720p, noting it's "a little bit more expensive" (approximately "8 cents per minute or per second" of video) but provides the highest quality.

Upon submission, the request goes into a queue, and the API returns a "request ID," which is crucial for tracking the job's status.

Monitoring and Retrieval of the Final Video

Since video generation is an asynchronous process that can take time, the workflow incorporates a polling mechanism:

  1. Wait Node: After submitting the face swap request, a Wait node is introduced, configured to pause the workflow for "1 minute" (or 2 minutes, depending on video length). This prevents continuous, rapid API calls.
  2. HTTP Request Node (Status Check): Another HTTP Request node is used to periodically check the status of the job using the previously obtained "request ID." The endpoint for this is https://api.file.ai/v1/models/wan-v2.2-14b-animate-replace/status/{request_id}.
  3. If Node: An If node checks if the status of the job is "completed."
    • If true (completed), the workflow proceeds to retrieve the final video.
    • If false (e.g., "in progress"), the workflow loops back to the Wait node, repeating the polling process.

The speaker notes that processing can take "between 15 to 20 minutes" for a 14-second video, advising users to "take a break" while it processes.

Once the status is "completed," a final HTTP Request node sends a GET request to retrieve the final video URL using the same status endpoint and the "request ID." This URL is then copied and pasted into a browser to download the generated face-swapped video.

Demonstration and Results

After the video processing completes (which took "a long time," as noted by the speaker), the final video URL is retrieved. The downloaded video is played, showcasing the face swap.

A side-by-side comparison of the original video and the AI-generated face-swapped video is presented. The speaker highlights the impressive realism:

  • "Even the head movements are very accurate."
  • The background "stays the same."
  • It "captures the lip movement, the lip sync, everything is great."
  • "Including the eye movements, everything is really aligned."

The speaker concludes that "this is a very powerful model" and demonstrates its high fidelity.

Conclusion and Future Outlook

The video reiterates the power and accessibility of AI for content creation. The speaker also promotes the "AI workshop community," which offers:

  • NADN for Beginners course: Comprehensive training on NADN automation.
  • Earn with NADN: Launch your AI Agency course: Guidance on monetizing AI skills and setting up an AI agency, emphasizing that "every business is going to require AI."
  • Voice course: Step-by-step training on selling voice agents, particularly in retail, with certification.
  • A supportive community of "close to 800 members from all over the world."

Looking ahead, the speaker plans to create a "full SaaS app" in a future video. This app would allow users to upload videos and images from a front-end interface, which would then send the data to the NADN backend for processing and video generation, further streamlining the user experience. The video concludes by encouraging viewers to like, subscribe, and stay tuned for more content on this and other AI models.

Chat with this Video

AI-Powered

Load the transcript when you're ready to chat so the initial page stays lighter.

Ready to summarize another video?

Summarize YouTube Video