SinoTechIntel Academic Portal
Open AccessDOI: 10.1631/ENG_ITEE_2026_0044Original Research

Three-dimensional affordance segmentation for object point cloud driven by language instructions

Jiaxuan DU¹,Hao WU¹,Qing MA¹,Guohui TIAN¹,Zhixian ZHAO¹,Shuwen LENG¹

School of Control Science and Engineering, Shandong University, Jinan 250061, China

Read Executive PreviewQuick FAQ
Three-dimensional affordance segmentation for object point cloud driven by language instructions
Graphical Abstract / Figure
Published In
Engineering Information Technology & Electronic Engineering
Published:August 19, 2025Edition:Vol. 32, Issue 8 • pp. 627-639Citation:Jiaxuan DU et al. (2025), Engineering Information Technology & Electronic Engineering
Impact Factor2.7 (Q2 - Springer)
Sponsored Research Partner
Keywords & Index Terms:visual affordancepoint cloud segmentationopen vocabularymultimodal fusionservice robotinstruction-driven6-DoF manipulationgrasp planning

Key Takeaways & Executive Findings

  • • Proposes a novel task of instruction-driven 3D object affordance segmentation from point clouds, bridging natural language and spatial manipulation reasoning. • Introduces the Instruction-Affordance Dataset (IAD) with 7,190 instances across 20 object categories and 624 manipulation instructions, including seen and unseen splits to test generalization. • Designs the IDAS network that integrates language instructions with point cloud features layer-by-layer, directly outputting task-relevant manipulation regions. • Demonstrates superior performance over existing methods under both seen and unseen instructions, showing strong generalization to novel commands and unknown affordances.
Sponsored Research Highlight

Abstract

The location where a robot grasps an object is closely related to the task type. For the same object, different user requirements may necessitate different grasping strategies. Visual affordance serves as a reliable source of prior knowledge for manipulation. Existing methods learn affordance from images or videos, but planar affordance lacks the spatial information required for 6-degree-of-freedom (6-DoF) manipulation. Furthermore, current approaches are limited to affordances associated with predefined categories and cannot directly infer affordances from user instructions. To address such limitations, we propose a novel task: instruction-driven three-dimensional (3D) object affordance segmentation. To support this research, we introduce an instruction–affordance dataset (IAD), a challenging dataset consisting of 7,190 object instances across 20 common object categories, paired with 624 manipulation instructions that specify the corresponding affordances. To evaluate generalization to novel commands, our dataset includes both seen and unseen settings. Building on this, we design an instruction-driven 3D affordance segmentation (IDAS) network, which extracts point cloud features and integrates instruction features layer by layer. Given a user instruction, our method segments suggested manipulation regions on the object’s point cloud, thereby guiding the selection of optimal grasp poses. Experimental results show that our method outperforms other related approaches under both seen and unseen settings, demonstrating generalization ability to diverse user commands and unknown affordances.

1. Introduction

Service robots require intelligence to accomplish tasks assigned by users (Wang et al., 2022). After actively approaching the target object (Liu SP et al., 2022), the robot needs to analyze how to manipulate it. Manipulation strategies for service robots must align with human expectations regarding how they interact with objects. For instance, when picking up a cup, humans generally prefer that the robot grasp the body of the cup instead of the rim, which is the part that comes into contact with their mouth. This preference reflects a more acceptable and hygienic approach to service. Additionally, the appropriate grasping location on an object often varies depending on user requirements. For example, when the task is to open a backpack, the robot needs to operate the zipper. However, if the task is to lift a backpack, it should instead grasp the handle.

Affordance acts as an informative prior that facilitates decision-making in robot service. Affordance refers to what the environment offers to an animal, including the opportunities it provides for interaction (Gibson, 1978). In the context of robot manipulation, affordances describe how objects present action possibilities for human interaction. These affordances are crucial for understanding how to manipulate objects effectively in robotics, and have been the subject of extensive study. Early research primarily focused on learning affordances from static images (Song et al., 2015; Roy and Todorovic, 2016; Do et al., 2018; Ardón et al., 2019) or from videos of human–object interaction (Fang K et al., 2018; Nagarajan et al., 2019; Goyal et al., 2022). More recent studies (Deng et al., 2021; Nguyen et al., 2023; Yang YH et al., 2023; Li YC et al., 2024) have advanced the concept by extending affordance segmentation into the three-dimensional (3D) domain using point cloud data. However, these affordance learning methods do not directly provide knowledge for manipulation; rather, they offer a broader conceptual understanding of the environment, which is difficult to apply in practice.

With the prompt of affordance, grasping approaches can provide more task-oriented and spatially reasonable grasp poses. Traditional grasp-planning approaches (Mousavian et al., 2019; Sundermeyer et al., 2021; Fang HS et al., 2023) primarily focus on physical feasibility, aiming to maximize the success rates of grasps without taking into account the humanlikeness or task relevance of the grip. As a result, these general-purpose methods are usually limited to simple manipulation tasks (Qin et al., 2023), such as sequentially picking up and placing objects on a tabletop. Therefore, learning grasp-relevant affordances is essential for handling diverse tasks.

SinoTechIntel Interactive Document Reader
Page 1–5 of Preview
100%
Download Full PDF

Loading authentic research manuscript (Pages 1–5)...

Sponsored Research Partner
Cite This Research Paper
Jiaxuan DU, Hao WU, Qing MA, Guohui TIAN, Zhixian ZHAO, Shuwen LENG (2025). Three-dimensional affordance segmentation for object point cloud driven by language instructions. Engineering Information Technology & Electronic Engineering. https://doi.org/10.1631/ENG_ITEE_2026_0044
SinoTechIntel Academic & Legal Disclaimer

Research & Educational Purpose Only:The translations, structured abstracts, analytical annotations, and data reports provided by SinoTechIntel are intended exclusively for academic research, internal corporate R&D, and educational benchmarking. They do not constitute formal engineering, chemical safety, legal, or professional advice.

Copyright & Intellectual Property Notice: Original copyright of the underlying source articles and experimental data remains with the respective authors, institutions, and original publishing journals. SinoTechIntel claims intellectual property only over its proprietary translations, analytical syntheses, and AEO structured enhancements in accordance with international fair use and academic citation principles.

Frequently Asked Questions

What is instruction-driven 3D object affordance segmentation?

It is a novel task that segments the regions of an object's 3D point cloud that are relevant to a given natural language instruction. This enables robots to manipulate objects according to user-specified tasks, such as grasping a cup by its body or operating a backpack zipper.

How does the IDAS network work?

The IDAS network extracts point cloud features and integrates instruction features layer by layer. Given a user instruction, it segments the suggested manipulation regions on the object's point cloud, which can then guide the selection of optimal grasp poses.

What is the Instruction-Affordance Dataset (IAD)?

IAD is a dataset introduced in this paper, containing 7,190 object instances across 20 common object categories and 624 manipulation instructions that specify corresponding affordances. It includes both seen and unseen settings to evaluate generalization to novel commands.

What are the main contributions of this research?

The main contributions are: (1) proposing the new task of instruction-driven 3D object affordance segmentation, (2) introducing the IAD dataset, (3) designing the IDAS network that integrates language and point cloud features, and (4) demonstrating strong generalization to unseen instructions and affordances.

How does this work benefit service robots?

It allows service robots to understand natural language instructions and map them to specific manipulation regions on objects. This leads to more human-like, task-relevant, and hygienic grasping behaviors, improving robot performance in diverse service scenarios.

Recommended Scientific Literature & Research Partners

Related Technical Papers & Translations

Research Paper
Design and optimization of a high-efficiency distillation process for cellulosic fuel ethanol integrated with thermal coupling and molecular sieve adsorption

Design and optimization of a high-efficiency distillation process for cellulosic fuel ethanol integrated with thermal coupling and molecular sieve adsorption

To address the challenges of high energy consumption and prominent costs in the traditional three-columns distillation process for cellulosic fuel ethanol, a distillation—molecular sieve coupling separation process is proposed. This process integrates a three-column (crude distillation column, first distillation column, second distillation column) system with a 3A molecular sieve adsorption deep dehydration unit. A thermal coupling network is constructed via differential pressure design (steam from medium/high-pressure columns as mutual heat sources, reboiler liquid waste heat for feed preheating), and molecular sieve adsorption conditions are optimized. The study first performs a thermodynamic consistency test on the ethanol—water system, determines optimal non-random two-liquid (NRTL) model binary interaction parameters via experimental data regression for Aspen Plus simulation. Aiming at minimum total annual cost (TAC), Aspen Plus is used to optimize process parameters (theoretical tray number, feed location, reflux ratio, side-draw position, etc.). Economic analysis shows this process reduces CO2 emission costs by 27.56%, TAC by 15.58% (to 5.123 × 106 USD·a-1), and increases ethanol purity to >99.6%, providing an effective solution for green, efficient separation.

Read Abstract & PDF
Research Paper
A cohesion loss model for determining residual strength of deep bedded sandstone

A cohesion loss model for determining residual strength of deep bedded sandstone

Rock residual strength, as an important input parameter, plays an indispensable role in proposing the reasonable and scientific scheme about stope design, underground tunnel excavation and stability evaluation of deep chambers. Therefore, previous residual strength models of rocks established were reviewed. And corresponding related problems were stated. Subsequently, starting from the effects of bedding and whole life-cycle evolution process, series of triaxial mechanical tests of deep bedded s

Read Abstract & PDF
Research Paper
Federated model with contrastive learning and adaptive control variates for human activity recognition

Federated model with contrastive learning and adaptive control variates for human activity recognition

Recent attention to privacy issues demands a communication-safe method for training human activity recognition (HAR) models on client activity data. Federated learning (FL) has become a compelling technique to facilitate model training between the server and clients while preserving data privacy. However, classical FL methods often assume independent and identically distributed (IID) data among clients. This assumption does not hold true in practical scenarios. Human activity in real-world scena

Read Abstract & PDF