17

2026-03

One Case a Day | China: Inventiveness Assessment of New AI Applications – "Image Processing Method" Case, Reexamination Decision No. 1694596 (2024)


Case Introduction

With the advancement of AI technology, AI is being integrated into an increasing number of fields, significantly enhancing efficiency. Meanwhile, practitioners are exploring ways to protect related innovations through patents. Regarding the examination criteria for whether new AI applications themselves possess inventiveness, relevant stakeholders need to understand and master them as soon as possible. If an invention merely involves replacing the recognition object of a known AI algorithm, its inventiveness may be limited. However, if a new application differs substantially from the prior art in terms of training data, input data, output data, model selection, etc., and addresses new technical problems, the related technical solution may be considered inventive.

Case Information

  • Application Number: 201810734681.2
  • Invention Title: Method and Apparatus for Processing Images
  • Reexamination Information: 1F397311
  • Reexamination Decision Information: No. 1694596
  • Decision Date: July 23, 2024

Key Points of the Decision

For inventions involving the application of artificial intelligence (AI) algorithms to specific scenarios, inventiveness should be evaluated by comprehensively considering the application scenario and the model algorithm, and assessing whether substantive adjustments or modifications have been made to the model algorithm features when applied to different scenarios. If the algorithm model of the involved patent, due to its application to different specific scenarios, differs substantially from the prior art in terms of training data, input data, output data, model selection, etc., and addresses different technical problems, the technical solution of the involved patent possesses inventiveness.

Case Summary

The relevant application Claim 1 protects a method for processing images, including:

  • Acquiring a captured image of a target object, wherein the target object includes a specified object and/or an object captured by the shooting device within a preset time period, and the target object includes a sports field;
  • Inputting the captured image into a pre-trained keypoint detection model corresponding to the target object to obtain a set of location information, wherein:
    • The location information includes location coordinates, which are used to represent the position of the keypoint indicated by the keypoint information in the keypoint information set in the captured image. The keypoint is a predetermined point on the target object, and different target objects correspond to different keypoint detection models. The keypoint detection model is trained based on captured images of the target object and the corresponding set of location information;
    • The location information also includes visibility information, which is used to represent the probability that the keypoint indicated by the keypoint information in the keypoint information set is displayed in the captured image.

For a specified or captured specific sports field object, through a pre-trained keypoint detection model and machine learning technology, the position and visibility screening of keypoints in the image are automatically determined, solving the problem of cumbersome manual annotation in the prior art and improving the efficiency and accuracy of image processing.

Prior Art

D1 (CN107679490A) is applied in the context of facial recognition. It uses a facial keypoint localization model to determine the coordinates of facial keypoints in any input facial image. By judging whether the facial part to which the keypoint coordinates belong is consistent with the facial part divided based on image content, the probability of occlusion of the facial image is determined, ultimately achieving quality evaluation of the facial image. The facial keypoint localization model is typically trained based on large-scale facial image datasets of different ages, races, and genders.

D2 (CN107590807A) is also applied in the context of facial recognition. It determines facial keypoint information in any input facial image and then evaluates the quality of the facial image based on this information. Although D2 uses a different method for determining facial keypoint information than D1—inputting the facial image into a convolutional neural network to obtain image feature information and parsing it to get the probability or coordinates of facial keypoint occlusion—its purpose remains the same.

Reexamination Decision

1. Compared with D1, the highlighted parts in Claim 1 are distinguishing features. The actual technical problem solved by Claim 1 is: how to determine the correspondence of the same target object in different captured images.

2. Claim 1 uses captured images and location information sets of each specific sports field object to train a keypoint detection model corresponding one-to-one with that specific sports field object. When multiple frames of images (such as images from different angles or at different times) of that specific sports field object are input into the corresponding keypoint detection model, the model outputs the different location coordinates and visibility information of the keypoints of that sports field object in different captured images. This allows for determining the geometric transformation relationship of the same sports field object's keypoints in different captured images, enabling game data analysis based on multiple frames of sports field images.

3. D1, although mentioning a keypoint localization model, only deals with facial images and does not involve sports field images. In D1, all facial images typically correspond to the same facial keypoint localization model, which is not one-to-one with specific target objects, and the facial recognition function in D1 does not require this. Therefore, a person skilled in the art cannot derive from D1 the technical motivation to set up a one-to-one corresponding keypoint localization model for each target object. The facial keypoint localization model in D1 only addresses keypoints and image quality in a single facial image and does not involve the correspondence of keypoints of the same target object across multiple images. Thus, D1 does not disclose the aforementioned distinguishing features and does not provide the motivation to adopt these distinguishing features to solve the technical problem of how to determine the correspondence of the same target object in different captured images.

4. D2 also does not involve the sports field captured images specified by the aforementioned distinguishing features and does not provide the technical motivation to establish a keypoint detection model one-to-one with the target object. Moreover, it does not face the technical problem of determining the correspondence of keypoints of the same sports field object in different captured images.

5. Additionally, although D2 mentions determining the probability of facial keypoint occlusion, its purpose is to evaluate facial image quality based on the degree of facial occlusion, which differs from the role of determining the probability of keypoints being displayed in the captured image in Claim 1 for screening visible keypoints. Therefore, D2 does not disclose the aforementioned distinguishing feature of "the probability of keypoints being displayed in the captured image." Thus, D2 does not disclose the aforementioned distinguishing features and does not provide any motivation to combine with D1 and apply it to sports field objects to solve the technical problem of how to determine the correspondence of the same target object in different captured images.

6. Therefore, neither D1 nor D2 discloses the aforementioned distinguishing features, nor do they provide corresponding technical motivation. Moreover, there is no evidence indicating that the aforementioned distinguishing features belong to common knowledge in the field. Through the application of these distinguishing features, Claim 1 achieves the beneficial effect of determining the correspondence of the same target object in different captured images.

Feng Shangjie (Gasoll Feng)

undefined

undefined