publications
Peer-reviewed publications and research papers in computer vision, machine learning, and 3D vision.
2026
-
Data Leakage Detection and De-duplication in Large Scale Geospatial Image DatasetsYeshwanth Kumar Adimoolam, Charalambos Poullis, and Melinos AverkiouIn Proc. CVPR, Jun 2026Selected for an oral presentation and as a CVPR 2026 award candidate.
This work audits large-scale geospatial image datasets for duplication and train-test leakage, and presents an efficient perceptual-hashing pipeline for detecting both problems before model training and evaluation.
-
EASE: Parametric Garment Design with Explicit and Local Ease ControlKristijan Bartol, Frieda Hentschel, Nataliya Sadretdinova, Benjamin Russig, Melinos Averkiou, Yordan Kyosev, and Stefan GumholdComputers & Graphics, Jun 2026EASE represents garment material allowance as an explicit, spatially varying design variable, enabling local editing, transfer across body shapes, and pose-aware adaptation.
2025
- Procedural Modelling of Aged Buildings: A Case Study of Mudbrick Houses in CyprusChristos Othonos, Gonzalo Besuievsky, Melinos Averkiou, Loizos Pelecanos, Gustavo Patow, and Yiorgos ChrysanthouACM Journal on Computing and Cultural Heritage, Dec 2025
This work presents a procedural system for simulating aging effects in buildings, using finite-element analysis to model degradation, weathering, and climate erosion in Cypriot mudbrick houses.
-
ShapeWords: Guiding Text-to-Image Synthesis with 3D Shape-Aware PromptsDmitry Petrov, Pradyumn Goyal, Divyansh Shivashok, Yuanming Tao, Melinos Averkiou, and Evangelos KalogerakisIn Proc. CVPR, Jun 2025ShapeWords embeds 3D shape information into specialized text tokens, enabling image synthesis that follows both a target geometry and a natural-language prompt.
-
Im2SurfTex: Surface Texture Generation via Neural Backprojection of Multi-View ImagesYiangos Georgiou, Marios Loizou, Melinos Averkiou, and Evangelos KalogerakisComputer Graphics Forum (Proc. SGP), 2025Im2SurfTex learns to combine multi-view diffusion-model outputs in texture space, improving continuity and reducing seams on complex 3D surfaces.
-
Pix2Poly: A Sequence Prediction Method for End-to-End Polygonal Building Footprint Extraction from Remote Sensing ImageryYeshwanth Kumar Adimoolam, Charalambos Poullis, and Melinos AverkiouIn Proc. WACV, Feb 2025Pix2Poly is an end-to-end transformer system that directly predicts explicit polygonal building footprints and road networks from remote-sensing imagery.
2024
-
GEM3D: GEnerative Medial Abstractions for 3D Shape SynthesisDmitry Petrov, Pradyumn Goyal, Vikas Thamizharasan, Vladimir G. Kim, Matheus Gadelha, Melinos Averkiou, Siddhartha Chaudhuri, and Evangelos KalogerakisIn Proc. SIGGRAPH, 2024We introduce GEM3D — a new deep, topology-aware generative model of 3D shapes. The key ingredient of our method is a neural skeleton-based representation encoding information on both shape topology and geometry. Through a denoising diffusion probabilistic model, our method first generates skeleton-based representations following the Medial Axis Transform (MAT), then generates surfaces through a skeleton-driven neural implicit formulation. The neural implicit takes into account the topological and geometric information stored in the generated skeleton representations to yield surfaces that are more topologically and geometrically accurate compared to previous neural field formulations. We discuss applications of our method in shape synthesis and point cloud reconstruction tasks, and evaluate our method both qualitatively and quantitatively. We demonstrate significantly more faithful surface reconstruction and diverse shape generation results compared to the state-of-the-art, also involving challenging scenarios of reconstructing and synthesizing structurally complex, high-genus shape surfaces from Thingi10K and ShapeNet.
-
FacadeNet: Conditional Facade Synthesis via Selective EditingYiangos Georgiou, Marios Loizou, Tom Kelly, and Melinos AverkiouIn Proc. WACV, 2024We introduce FacadeNet, a deep learning approach for synthesizing building facade images from diverse viewpoints. Our method employs a conditional GAN, taking a single view of a facade along with the desired viewpoint information and generates an image of the facade from the distinct viewpoint. To precisely modify view-dependent elements like windows and doors while preserving the structure of view-independent components such as walls, we introduce a selective editing module. This module leverages image embeddings extracted from a pretrained vision transformer. Our experiments demonstrated state-of-the-art performance on building facade generation, surpassing alternative methods
2023
-
Cross-Shape Attention for Part Segmentation of 3D Point CloudsM. Loizou, S. Garg, D. Petrov, M. Averkiou, and E. KalogerakisComputer Graphics Forum (Proc. SGP), 2023We present a deep learning method that propagates point-wise feature representations across shapes within a collection for the purpose of 3D shape segmentation. We propose a cross-shape attention mechanism to enable interactions between a shape’s point-wise features and those of other shapes. The mechanism assesses both the degree of interaction between points and also mediates feature propagation across shapes, improving the accuracy and consistency of the resulting point-wise feature representations for shape segmentation. Our method also proposes a shape retrieval measure to select suitable shapes for cross-shape attention operations for each test shape. Our experiments demonstrate that our approach yields state-of-the-art results in the popular PartNet dataset.
-
An artificial neural network framework for classifying the style of cypriot hybrid examples of built heritage in 3DG. Artopoulos, M. I. Maslioukova, C. Zavou, M. Loizou, M. Deligiorgi, and M. AverkiouJournal of Cultural Heritage, 2023The article presents a workflow based on Deep Neural Networks (DNNs) and Support Vector Machine (SVM) for identifying architectural stylistic influences of segmented building parts of Cypriot historical architecture in 3D. The research contributes in the field of Digital Cultural Heritage (DCH) by applying Machine Learning (ML) and Deep Learning (DL) on recently published DCH data, with the aim to accelerate the segmentation and annotation process of Historic Building Information modelling (HBIM) that is currently based on time-consuming manual processes. The method presented works on reality captured data by 3D documentation techniques, precisely, Terrestrial Laser Scanning (TLS) or Photogrammetry. This workflow was developed to enable the operation of an online platform,11https://annfass-srv.cs.ucy.ac.cy. which also provides access to the building data presented here. Ultimately, the results of the presented method are accessible to scholars and students via this platform which provides multiple functionalities for researchers in the field to use.
2021
-
A 3D digitisation workflow for architecture-specific annotation of built heritageMarissia Deligiorgi, Maria I. Maslioukova, Melinos Averkiou, Andreas C. Andreou, Pratheba Selvaraju, Evangelos Kalogerakis, Gustavo Patow, Yiorgos Chrysanthou, and George ArtopoulosJournal of Archaeological Science: Reports, 2021Contemporary discourse points to the central role that heritage plays in the process of enabling groups of various cultural or ethnic background to strengthen their feeling of belonging and sharing in society. Safeguarding heritage is also valued highly in the priorities of the European Commission. As a result, there have been several long-term initiatives involving the digitisation, annotation and cataloguing of tangible cultural heritage in museums and collections. Specifically, for built heritage, a pressing challenge is that historical monuments such as buildings, temples, churches or city fortification infrastructures are hard to document due to their historic palimpsest; spatial transformations, actions of destruction, reuse of material, or continuous urban development that covers traces and changes the formal integrity and identity of a cultural heritage site. The ability to reason about a monument’s form is crucial for efficient documentation and cataloguing. This paper presents a 3D digitisation workflow through the involvement of reality capture technologies for the annotation and structure analysis of built heritage with the use of 3D Convolutional Neural Networks (3D CNNs) for classification purposes. The presented workflow contributes a new approach to the identification of a building’s architectural components (e.g., arch, dome) and to the study of the stylistic influences (e.g., Gothic, Byzantine) of building parts. In doing so this workflow can assist in tracking a building’s history, identifying its construction period and comparing it to other buildings of the same period. This process can contribute to educational and research activities, as well as facilitate the automated classification of datasets in digital repositories for scholarly research in digital humanities.
-
Projective Urban TexturingYiangos Georgiou, Melinos Averkiou, Tom Kelly, and Evangelos KalogerakisIn Proc. 3DV, Dec 2021This paper proposes a method for automatic generation of textures for 3D city meshes in immersive urban environments. Many recent pipelines capture or synthesize large quantities of city geometry using scanners or procedural modeling pipelines. Such geometry is intricate and realistic, however the generation of photo-realistic textures for such large scenes remains a problem. We propose to generate textures for input target 3D meshes driven by the textural style present in readily available datasets of panoramic photos capturing urban environments. Re-targeting such 2D datasets to 3D geometry is challenging because the underlying shape, size, and layout of the urban structures in the photos do not correspond to the ones in the target meshes. Photos also often have objects (e.g., trees, vehicles) that may not even be present in the target geometry. To address these issues we present a method, called Projective Urban Texturing (PUT), which re-targets textural style from real-world panoramic images to unseen urban meshes. PUT relies on contrastive and adversarial training of a neural architecture designed for unpaired image-to-texture translation. The generated textures are stored in a texture atlas applied to the target 3D mesh geometry. To promote texture consistency, PUT employs an iterative procedure in which texture synthesis is conditioned on previously generated, adjacent textures. We demonstrate both quantitative and qualitative evaluation of the generated textures.
-
BuildingNet: Learning to Label 3D BuildingsP. Selvaraju, M. Nabail, M. Loizou, M. Maslioukova, M. Averkiou, A. Andreou, S. Chaudhuri, and E. KalogerakisIn Proc. ICCV, 2021We introduce BuildingNet: (a) a large-scale dataset of 3D building models whose exteriors are consistently labeled, and (b) a graph neural network that labels building meshes by analyzing spatial and structural relations of their geometric primitives. To create our dataset, we used crowdsourcing combined with expert guidance, resulting in 513K annotated mesh primitives, grouped into 292K semantic part components across 2K building models. The dataset covers several building categories, such as houses, churches, skyscrapers, town halls, libraries, and castles. We include a benchmark for evaluating mesh and point cloud labeling. Buildings have more challenging structural complexity compared to objects in existing benchmarks (e.g., ShapeNet, PartNet), thus, we hope that our dataset can nurture the development of algorithms that are able to cope with such large-scale geometric data for both vision and graphics tasks e.g., 3D semantic segmentation, part-based generative models, correspondences, texturing, and analysis of point cloud data acquired from real-world buildings. Finally, we show that our mesh-based graph neural network significantly improves performance over several baselines for labeling 3D meshes. Our project page www.buildingnet.org includes our dataset and code.
2020
-
Learning Part Boundaries from 3D Point CloudsMarios Loizou, Melinos Averkiou, and Evangelos KalogerakisComputer Graphics Forum (Proc. SGP), 2020We present a method that detects boundaries of parts in 3D shapes represented as point clouds. Our method is based on a graph convolutional network architecture that outputs a probability for a point to lie in an area that separates two or more parts in a 3D shape. Our boundary detector is quite generic: it can be trained to localize boundaries of semantic parts or geometric primitives commonly used in 3D modeling. Our experiments demonstrate that our method can extract more accurate boundaries that are closer to ground-truth ones compared to alternatives. We also demonstrate an application of our network to fine-grained semantic shape segmentation, where we also show improvements in terms of part labeling performance.
2018
-
Learning Material-Aware Local Descriptors for 3D ShapesHubert Lin, Melinos Averkiou, Evangelos Kalogerakis, Balazs Kovacs, Siddhant Ranade, Vladimir Kim, Siddhartha Chaudhuri, and Kavita BalaIn Proc. 3DV, Sep 2018Material understanding is critical for design, geometric modeling, and analysis of functional objects. We enable material-aware 3D shape analysis by employing a projective convolutional neural network architecture to learn material-aware descriptors from view-based representations of 3D points for point-wise material classification or material-aware retrieval. Unfortunately, only a small fraction of shapes in 3D repositories are labeled with physical materials, posing a challenge for learning methods. To address this challenge, we crowdsource a dataset of 3080 3D shapes with part-wise material labels. We focus on furniture models which exhibit interesting structure and material variability. In addition, we also contribute a high-quality expert-labeled benchmark of 115 shapes from Herman-Miller and IKEA for evaluation. We further apply a mesh-aware conditional random field, which incorporates rotational and reflective symmetries, to smooth our local material predictions across neighboring surface patches. We demonstrate the effectiveness of our learned descriptors for automatic texturing, material-aware retrieval, and physical simulation.
2017
-
3D Shape Segmentation with Projective Convolutional NetworksEvangelos Kalogerakis, Melinos Averkiou, Subhransu Maji, and Siddhartha ChaudhuriIn Proc. CVPR, Jul 2017This paper introduces a deep architecture for segmenting 3D objects into their labeled semantic parts. Our architecture combines image-based Fully Convolutional Networks (FCNs) and surface-based Conditional Random Fields (CRFs) to yield coherent segmentations of 3D shapes. The image-based FCNs are used for efficient view-based reasoning about 3D object parts. Through a special projection layer, FCN outputs are effectively aggregated across multiple views and scales, then are projected onto the 3D object surfaces. Finally, a surface-based CRF combines the projected outputs with geometric consistency cues to yield coherent segmentations. The whole architecture (multi-view FCNs and CRF) is trained end-to-end. Our approach significantly outperforms the existing state-of-the-art methods in the currently largest segmentation benchmark (ShapeNet). Finally, we demonstrate promising segmentation results on noisy 3D shapes acquired from consumer-grade depth cameras.
-
Co-Locating Style-Defining Elements on 3D ShapesRuizhen Hu, Wenchao Li, Oliver Van Kaick, Hui Huang, Melinos Averkiou, Daniel Cohen-Or, and Hao ZhangACM TOG, Jun 2017We introduce a method for co-locating style-defining elements over a set of 3D shapes. Our goal is to translate high-level style descriptions, such as “Ming” or “European” for furniture models, into explicit and localized regions over the geometric models that characterize each style. For each style, the set of style-defining elements is defined as the union of all the elements that are able to discriminate the style. Another property of the style-defining elements is that they are frequently occurring, reflecting shape characteristics that appear across multiple shapes of the same style. Given an input set of 3D shapes spanning multiple categories and styles, where the shapes are grouped according to their style labels, we perform a cross-category co-analysis of the shape set to learn and spatially locate a set of defining elements for each style. This is accomplished by first sampling a large number of candidate geometric elements and then iteratively applying feature selection to the candidates, to extract style-discriminating elements until no additional elements can be found. Thus, for each style label, we obtain sets of discriminative elements that together form the superset of defining elements for the style. We demonstrate that the co-location of style-defining elements allows us to solve problems such as style classification, and enables a variety of applications such as style-revealing view selection, style-aware sampling, and style-driven modeling for 3D shapes.
2016
-
Autocorrelation Descriptor for Efficient Co-Alignment of 3D Shape CollectionsMelinos Averkiou, Vladimir G. Kim, and Niloy J. MitraComputer Graphics Forum, 2016Abstract Co-aligning a collection of shapes to a consistent pose is a common problem in shape analysis with applications in shape matching, retrieval and visualization. We observe that resolving among some orientations is easier than others, for example, a common mistake for bicycles is to align front-to-back, while even the simplest algorithm would not erroneously pick orthogonal alignment. The key idea of our work is to analyse rotational autocorrelations of shapes to facilitate shape co-alignment. In particular, we use such an autocorrelation measure of individual shapes to decide which shape pairs might have well-matching orientations; and, if so, which configurations are likely to produce better alignments. This significantly prunes the number of alignments to be examined, and leads to an efficient, scalable algorithm that performs comparably to state-of-the-art techniques on benchmark data sets, but requires significantly fewer computations, resulting in 2–16× speed improvement in our tests.
2015
-
Data-driven Modelling of Shape StructureMelinos AverkiouUniversity College London, Aug 2015In recent years, the study of shape structure has shown great promise, by taking steps towards exposing shape semantics and functionality to algorithms spanning a wide range of areas in computer graphics and vision. By shape structure, we refer to the set of parts that make a shape, the relations between these parts, and the ways in which they correspond and vary between shapes of the same family. These developments have been largely driven by the abundance of 3D data, with collections of 3D models becoming increasingly prominent and websites such as Trimble 3D Warehouse offering millions of free 3D models to the public. The ability to use large amounts of data inside these shape collections for discovering shape structure has made novel approaches to acquisition, modelling, fabrication, and recognition of 3D objects possible. Discovering and modelling the structure of shapes using such data is therefore of great importance. In this thesis we address the problem of discovering and modelling shape structure from large, diverse and unorganized shape collections. Our hypothesis is that by using the large amounts of data inside such shape collections we can discover and model shape structure, and thus use such information to enable structure-aware tools for 3D modelling, including shape exploration, synthesis and editing. We make three key contributions. First, we propose an efficient algorithm for co-aligning large and diverse collections of shapes, to tackle the first challenge in detecting shape structure, which is to place shapes in a common coordinate frame. Then, we introduce a method to parameterize shapes in terms of locations and sizes of their parts, and we demonstrate its application to concurrently exploring a shape collection and synthesizing new shapes. Finally, we define a meta-representation for a shape family, which models the relations of shape parts to capture the main geometric characteristics of the family, and we demonstrate how it can be used to explore shape collections and intelligently edit shapes.
-
Object Proposals Estimation in Depth Image Using Compact 3D Shape ManifoldsShuai Zheng, Victor Adrian Prisacariu, Melinos Averkiou, Ming-Ming Cheng, Niloy J. Mitra, Jamie Shotton, Philip H. S. Torr, and Carsten RotherIn German Conference on Pattern Recognition, 2015Man-made objects, such as chairs, often have very large shape variations, making it challenging to detect them. In this work we investigate the task of finding particular object shapes from a single depth image. We tackle this task by exploiting the inherently low dimensionality in the object shape variations, which we discover and encode as a compact shape space. Starting from any collection of 3D models, we first train a low dimensional Gaussian Process Latent Variable Shape Space. We then sample this space, effectively producing infinite amounts of shape variations, which are used for training. Additionally, to support fast and accurate inference, we improve the standard 3D object category proposal generation pipeline by applying a shallow convolutional neural network-based filtering stage. This combination leads to considerable improvements for proposal generation, in both speed and accuracy. We compare our full system to previous state-of-the-art approaches, on four different shape classes, and show a clear improvement.
2014
-
Meta-Representation of Shape FamiliesNoa Fish, Melinos Averkiou, Oliver Kaick, Olga Sorkine-Hornung, Daniel Cohen-Or, and Niloy J. MitraACM TOG (Proc. SIGGRAPH), Jul 2014We introduce a meta-representation that represents the essence of a family of shapes. The meta-representation learns the configurations of shape parts that are common across the family, and encapsulates this knowledge with a system of geometric distributions that encode relative arrangements of parts. Thus, instead of predefined priors, what characterizes a shape family is directly learned from the set of input shapes. The meta-representation is constructed from a set of co-segmented shapes with known correspondence. It can then be used in several applications where we seek to preserve the identity of the shapes as members of the family. We demonstrate applications of the meta-representation in exploration of shape repositories, where interesting shape configurations can be examined in the set; guided editing, where models can be edited while maintaining their familial traits; and coupled editing, where several shapes can be collectively deformed by directly manipulating the distributions in the meta-representation. We evaluate the efficacy of the proposed representation on a variety of shape collections.
-
Recurring part arrangements in shape collectionsYouyi Zheng, Daniel Cohen-Or, Melinos Averkiou, and Niloy J. MitraComputer Graphics Forum (Proc. Eurographics), 2014Abstract Extracting semantically related parts across models remains challenging, especially without supervision. The common approach is to co-analyze a model collection, while assuming the existence of descriptive geometric features that can directly identify related parts. In the presence of large shape variations, common geometric features, however, are no longer sufficiently descriptive. In this paper, we explore an indirect top-down approach, where instead of part geometry, part arrangements extracted from each model are compared. The key observation is that while a direct comparison of part geometry can be ambiguous, part arrangements, being higher level structures, remain consistent, and hence can be used to discover latent commonalities among semantically related shapes. We show that our indirect analysis leads to the detection of recurring arrangements of parts, which are otherwise difficult to discover in a direct unsupervised setting. We evaluate our algorithm on ground truth datasets and report advantages over geometric similarity-based bottom-up co-segmentation algorithms.
-
ShapeSynth: Parameterizing model collections for coupled shape exploration and synthesisMelinos Averkiou, Vladimir G. Kim, Youyi Zheng, and Niloy J. MitraComputer Graphics Forum (Proc. Eurographics), 2014Abstract Recent advances in modeling tools enable non-expert users to synthesize novel shapes by assembling parts extracted from model databases. A major challenge for these tools is to provide users with relevant parts, which is especially difficult for large repositories with significant geometric variations. In this paper we analyze unorganized collections of 3D models to facilitate explorative shape synthesis by providing high-level feedback of possible synthesizable shapes. By jointly analyzing arrangements and shapes of parts across models, we hierarchically embed the models into low-dimensional spaces. The user can then use the parameterization to explore the existing models by clicking in different areas or by selecting groups to zoom on specific shape clusters. More importantly, any point in the embedded space can be lifted to an arrangement of parts to provide an abstracted view of possible shape variations. The abstraction can further be realized by appropriately deforming parts from neighboring models to produce synthesized geometry. Our experiments show that users can rapidly generate plausible and diverse shapes using our system, which also performs favorably with respect to previous modeling tools.
2011
-
Comparison of relative (mouse-like) and absolute (tablet-like) interaction with a large stereoscopic workspaceMelinos Averkiou and Neil A. DodgsonIn Stereoscopic Displays and Applications XXII, 2011We compare two different modes of interaction with a large stereoscopic display, where the physical pointing device is in a volume distinct from the display volume. In absolute mode, the physical pointer’s position exactly maps to the virtual pointer’s position in the display volume, analogous to a 2D graphics table and 2D screen. In relative mode, the connection between the physical pointer’s motion and the motion of the virtual pointer in the display volume is analogous to that obtained with a 2D mouse and 2D screen. Both statistical analysis and participants’ feedback indicated a strong preference for absolute mode over relative mode. This is in contrast to 2D displays where relative mode (mouse) is far more prevalent than absolute mode (tablet). We also compared head-tracking against no head-tracking. There was no statistically-significant advantage to using head-tracking, however almost all participants strongly favoured head-tracking.