Skip to content

#Image generation

0 today

Sep 30

Wednesday
  1. How Diffusion Controller unifies and simplifies AI image generation

    Google Research proposed Diffusion Controller, reframing diffusion-model denoising as a continuous control problem. A lightweight steering-damper network dynamically adjusts generation trajectories while the base model remains frozen. Evaluated on Stable Diffusion v1.4 using HPS-v2, its fully unlocked version achieved a 90% win rate against the baseline, and it supports customised control of closed models without access to internal weights.

Sep 29

Tuesday
  1. Generate images and video with vLLM-Omni on SageMaker AI – Part 2

    AWS published a tutorial deploying two endpoints from the same vLLM-Omni DLC image on SageMaker AI: a real-time endpoint running FLUX.2-klein-4B for images and an asynchronous endpoint running Wan2.1-VACE-1.3B for image-conditioned video. A text prompt first generates a PNG, which is sent with a motion prompt through Amazon S3 to the video endpoint; the resulting MP4 is retrieved from S3. The example includes a command-line workflow and an optional Streamlit interface.

Sep 8

Tuesday

Jul 23

Thursday

Jul 16

Thursday

Jul 1

Wednesday

Apr 23

Thursday
  1. It's all about the angle: Your photos, re-composed

    Google Research introduced a new method in Google Photos’ Auto frame to change camera viewpoints after a photo is taken. An internal 3D point-map estimation model reconstructs the scene and focal length, then a generative latent diffusion model fills gaps exposed by the new view. It automatically detects faces and wide-angle distortion to suggest ideal framing. The feature already applies automatically to photos containing people, with reframed versions available among Auto frame candidates.

Jan 10

Saturday
  1. Information-Driven Design of Imaging Systems

    Berkeley AI Research proposed a mutual-information framework to evaluate and optimise imaging from noisy measurements and noise models. A NeurIPS 2025 paper validates decoder performance predictions for colour photography, radio astronomy, lensless imaging and microscopy. Designs match end-to-end state-of-the-art approaches with less memory and compute, without task-specific decoder design.