fixing issue on propagation organization. added --config to workflow definition. added oozie_app to communtiy project

mergin with branch beta
Merge pull request 'FIX: GroupEntitiesSparkJob deletes whole graph outputPath instead of its temporary folder' (#351 ) from fix_consistency_missing_rels into beta
2023-10-20 10:14:20 +02:00 · 2023-10-19 09:04:35 +02:00 · 2023-10-17 08:40:23 +02:00 · 2023-10-17 08:40:07 +02:00 · 2023-10-17 07:54:01 +02:00 · 2023-10-16 11:56:18 +02:00
209 changed files with 7227 additions and 2623 deletions
--- a/.gitignore
+++ b/.gitignore
@ -26,3 +26,4 @@ spark-warehouse
 /**/*.log
 /**/.factorypath
 /**/.scalafmt.conf
+/.java-version
--- a/README.md
+++ b/README.md
@ -1,2 +1,128 @@
 # dnet-hadoop
-Dnet-hadoop is the project that defined all the OOZIE workflows for the OpenAIRE Graph construction, processing, provisioning.
+
+Dnet-hadoop is the project that defined all the [OOZIE workflows](https://oozie.apache.org/) for the OpenAIRE Graph construction, processing, provisioning.
+
+How to build, package and run oozie workflows
+====================
+
+Oozie-installer is a utility allowing building, uploading and running oozie workflows. In practice, it creates a `*.tar.gz` 
+package that contains resources that define a workflow and some helper scripts.
+
+This module is automatically executed when running:
+
+`mvn package -Poozie-package -Dworkflow.source.dir=classpath/to/parent/directory/of/oozie_app`
+
+on module having set:
+
+```
+<parent>
+    <groupId>eu.dnetlib.dhp</groupId>
+    <artifactId>dhp-workflows</artifactId>
+</parent>
+```
+
+in `pom.xml` file. `oozie-package` profile initializes oozie workflow packaging, `workflow.source.dir` property points to 
+a workflow (notice: this is not a relative path but a classpath to directory usually holding `oozie_app` subdirectory).
+
+The outcome of this packaging is `oozie-package.tar.gz` file containing inside all the resources required to run Oozie workflow:
+
+- jar packages
+- workflow definitions
+- job properties
+- maintenance scripts
+
+Required properties
+====================
+
+In order to include proper workflow within package, `workflow.source.dir` property has to be set. It could be provided 
+by setting `-Dworkflow.source.dir=some/job/dir` maven parameter.
+
+In oder to define full set of cluster environment properties one should create `~/.dhp/application.properties` file with 
+the following properties:
+
+- `dhp.hadoop.frontend.user.name` - your user name on hadoop cluster and frontend machine
+- `dhp.hadoop.frontend.host.name` - frontend host name
+- `dhp.hadoop.frontend.temp.dir` - frontend directory for temporary files
+- `dhp.hadoop.frontend.port.ssh` - frontend machine ssh port
+- `oozieServiceLoc` - oozie service location required by run_workflow.sh script executing oozie job
+- `nameNode` - name node address
+- `jobTracker` - job tracker address
+- `oozie.execution.log.file.location` - location of file that will be created when executing oozie job, it contains output 
+produced by `run_workflow.sh` script (needed to obtain oozie job id)
+- `maven.executable` - mvn command location, requires parameterization due to a different setup of CI cluster
+- `sparkDriverMemory` - amount of memory assigned to spark jobs driver
+- `sparkExecutorMemory` - amount of memory assigned to spark jobs executors
+- `sparkExecutorCores` - number of cores assigned to spark jobs executors
+
+All values will be overriden with the ones from `job.properties` and eventually `job-override.properties` stored in module's 
+main folder.
+
+When overriding properties from `job.properties`, `job-override.properties` file can be created in main module directory 
+(the one containing `pom.xml` file) and define all new properties which will override existing properties. 
+One can provide those properties one by one as command line `-D` arguments.
+
+Properties overriding order is the following:
+
+1. `pom.xml` defined properties (located in the project root dir)
+2. `~/.dhp/application.properties` defined properties
+3. `${workflow.source.dir}/job.properties`
+4. `job-override.properties` (located in the project root dir)
+5. `maven -Dparam=value`
+
+where the maven `-Dparam` property is overriding all the other ones.
+
+Workflow definition requirements
+====================
+
+`workflow.source.dir` property should point to the following directory structure:
+
+	[${workflow.source.dir}]
+		|
+		|-job.properties (optional)
+		|
+		\-[oozie_app]
+			|
+			\-workflow.xml
+
+This property can be set using maven `-D` switch.
+
+`[oozie_app]` is the default directory name however it can be set to any value as soon as `oozieAppDir` property is 
+provided with directory name as value.
+
+Sub-workflows are supported as well and sub-workflow directories should be nested within `[oozie_app]` directory.
+
+Creating oozie installer step-by-step
+=====================================
+
+Automated oozie-installer steps are the following:
+
+1. creating jar packages:  `*.jar` and `*tests.jar` along with copying all dependencies in `target/dependencies`
+2. reading properties from maven, `~/.dhp/application.properties`, `job.properties`, `job-override.properties`
+3. invoking priming mechanism linking resources from import.txt file (currently resolving subworkflow resources)
+4. assembling shell scripts for preparing Hadoop filesystem, uploading Oozie application and starting workflow
+5. copying whole `${workflow.source.dir}` content to `target/${oozie.package.file.name}`
+6. generating updated `job.properties` file in `target/${oozie.package.file.name}` based on maven, 
+`~/.dhp/application.properties`, `job.properties` and `job-override.properties`
+7. creating `lib` directory (or multiple directories for sub-workflows for each nested directory) and copying jar packages 
+created at step (1) to each one of them
+8. bundling whole `${oozie.package.file.name}` directory into single tar.gz package
+
+Uploading oozie package and running workflow on cluster
+=======================================================
+
+In order to simplify deployment and execution process two dedicated profiles were introduced:
+
+- `deploy`
+- `run`
+
+to be used along with `oozie-package` profile e.g. by providing `-Poozie-package,deploy,run` maven parameters.
+
+The `deploy` profile supplements packaging process with:
+1) uploading oozie-package via scp to `/home/${user.name}/oozie-packages` directory on `${dhp.hadoop.frontend.host.name}` machine
+2) extracting uploaded package
+3) uploading oozie content to hadoop cluster HDFS location defined in `oozie.wf.application.path` property (generated dynamically by maven build process, based on `${dhp.hadoop.frontend.user.name}` and `workflow.source.dir` properties)
+
+The `run` profile introduces:
+1) executing oozie application uploaded to HDFS cluster using `deploy` command. Triggers `run_workflow.sh` script providing runtime properties defined in `job.properties` file.
+
+Notice: ssh access to frontend machine has to be configured on system level and it is preferable to set key-based authentication in order to simplify remote operations.
--- a/dhp-common/src/main/java/eu/dnetlib/dhp/common/Constants.java
+++ b/dhp-common/src/main/java/eu/dnetlib/dhp/common/Constants.java
@ -51,6 +51,7 @@ public class Constants {
 	public static final String RETRY_DELAY = "retryDelay";
 	public static final String CONNECT_TIMEOUT = "connectTimeOut";
 	public static final String READ_TIMEOUT = "readTimeOut";
+	public static final String REQUEST_METHOD = "requestMethod";
 	public static final String FROM_DATE_OVERRIDE = "fromDateOverride";
 	public static final String UNTIL_DATE_OVERRIDE = "untilDateOverride";

--- a/dhp-common/src/main/java/eu/dnetlib/dhp/common/collection/HttpClientParams.java
+++ b/dhp-common/src/main/java/eu/dnetlib/dhp/common/collection/HttpClientParams.java
@ -1,6 +1,9 @@

 package eu.dnetlib.dhp.common.collection;

+import java.util.HashMap;
+import java.util.Map;
+
 /**
 * Bundles the http connection parameters driving the client behaviour.
 */
@ -13,6 +16,8 @@ public class HttpClientParams {
 	public static int _connectTimeOut = 10; // seconds
 	public static int _readTimeOut = 30; // seconds

+	public static String _requestMethod = "GET";
+
 	/**
 	 * Maximum number of allowed retires before failing
 	 */
@ -38,17 +43,30 @@ public class HttpClientParams {
 	 */
 	private int readTimeOut;

+	/**
+	 * Custom http headers
+	 */
+	private Map<String, String> headers;
+
+	/**
+	 * Request method (i.e., GET, POST etc)
+	 */
+	private String requestMethod;
+
 	public HttpClientParams() {
-		this(_maxNumberOfRetry, _requestDelay, _retryDelay, _connectTimeOut, _readTimeOut);
+		this(_maxNumberOfRetry, _requestDelay, _retryDelay, _connectTimeOut, _readTimeOut, new HashMap<>(),
+			_requestMethod);
 	}

 	public HttpClientParams(int maxNumberOfRetry, int requestDelay, int retryDelay, int connectTimeOut,
-		int readTimeOut) {
+		int readTimeOut, Map<String, String> headers, String requestMethod) {
 		this.maxNumberOfRetry = maxNumberOfRetry;
 		this.requestDelay = requestDelay;
 		this.retryDelay = retryDelay;
 		this.connectTimeOut = connectTimeOut;
 		this.readTimeOut = readTimeOut;
+		this.headers = headers;
+		this.requestMethod = requestMethod;
 	}

 	public int getMaxNumberOfRetry() {
@ -91,4 +109,19 @@ public class HttpClientParams {
 		this.readTimeOut = readTimeOut;
 	}

+	public Map<String, String> getHeaders() {
+		return headers;
+	}
+
+	public void setHeaders(Map<String, String> headers) {
+		this.headers = headers;
+	}
+
+	public String getRequestMethod() {
+		return requestMethod;
+	}
+
+	public void setRequestMethod(String requestMethod) {
+		this.requestMethod = requestMethod;
+	}
 }
--- a/dhp-common/src/main/java/eu/dnetlib/dhp/common/collection/HttpConnector2.java
+++ b/dhp-common/src/main/java/eu/dnetlib/dhp/common/collection/HttpConnector2.java
@ -107,7 +107,14 @@ public class HttpConnector2 {
 			urlConn.setReadTimeout(getClientParams().getReadTimeOut() * 1000);
 			urlConn.setConnectTimeout(getClientParams().getConnectTimeOut() * 1000);
 			urlConn.addRequestProperty(HttpHeaders.USER_AGENT, userAgent);
+			urlConn.setRequestMethod(getClientParams().getRequestMethod());

+			// if provided, add custom headers
+			if (!getClientParams().getHeaders().isEmpty()) {
+				for (Map.Entry<String, String> headerEntry : getClientParams().getHeaders().entrySet()) {
+					urlConn.addRequestProperty(headerEntry.getKey(), headerEntry.getValue());
+				}
+			}
 			if (log.isDebugEnabled()) {
 				logHeaderFields(urlConn);
 			}
--- a/dhp-common/src/main/java/eu/dnetlib/dhp/oa/merge/DispatchEntitiesSparkJob.java
+++ b/dhp-common/src/main/java/eu/dnetlib/dhp/oa/merge/DispatchEntitiesSparkJob.java
@ -1,98 +0,0 @@
-
-package eu.dnetlib.dhp.oa.merge;
-
-import static eu.dnetlib.dhp.common.SparkSessionSupport.runWithSparkSession;
-
-import java.util.Objects;
-import java.util.Optional;
-
-import org.apache.commons.io.IOUtils;
-import org.apache.commons.lang3.StringUtils;
-import org.apache.spark.SparkConf;
-import org.apache.spark.api.java.function.FilterFunction;
-import org.apache.spark.api.java.function.MapFunction;
-import org.apache.spark.sql.*;
-import org.slf4j.Logger;
-import org.slf4j.LoggerFactory;
-
-import eu.dnetlib.dhp.application.ArgumentApplicationParser;
-import eu.dnetlib.dhp.common.HdfsSupport;
-import eu.dnetlib.dhp.schema.common.ModelSupport;
-
-public class DispatchEntitiesSparkJob {
-
-	private static final Logger log = LoggerFactory.getLogger(DispatchEntitiesSparkJob.class);
-
-	public static void main(String[] args) throws Exception {
-
-		String jsonConfiguration = IOUtils
-			.toString(
-				Objects
-					.requireNonNull(
-						DispatchEntitiesSparkJob.class
-							.getResourceAsStream(
-								"/eu/dnetlib/dhp/oa/merge/dispatch_entities_parameters.json")));
-		final ArgumentApplicationParser parser = new ArgumentApplicationParser(jsonConfiguration);
-		parser.parseArgument(args);
-
-		Boolean isSparkSessionManaged = Optional
-			.ofNullable(parser.get("isSparkSessionManaged"))
-			.map(Boolean::valueOf)
-			.orElse(Boolean.TRUE);
-		log.info("isSparkSessionManaged: {}", isSparkSessionManaged);
-
-		String inputPath = parser.get("inputPath");
-		log.info("inputPath: {}", inputPath);
-
-		String outputPath = parser.get("outputPath");
-		log.info("outputPath: {}", outputPath);
-
-		boolean filterInvisible = Boolean.valueOf(parser.get("filterInvisible"));
-		log.info("filterInvisible: {}", filterInvisible);
-
-		SparkConf conf = new SparkConf();
-		runWithSparkSession(
-			conf,
-			isSparkSessionManaged,
-			spark -> {
-				HdfsSupport.remove(outputPath, spark.sparkContext().hadoopConfiguration());
-				dispatchEntities(spark, inputPath, outputPath, filterInvisible);
-			});
-	}
-
-	private static void dispatchEntities(
-		SparkSession spark,
-		String inputPath,
-		String outputPath,
-		boolean filterInvisible) {
-
-		Dataset<String> df = spark.read().textFile(inputPath);
-
-		ModelSupport.oafTypes.entrySet().parallelStream().forEach(entry -> {
-			String entityType = entry.getKey();
-			Class<?> clazz = entry.getValue();
-
-			if (!entityType.equalsIgnoreCase("relation")) {
-				Dataset<Row> entityDF = spark
-					.read()
-					.schema(Encoders.bean(clazz).schema())
-					.json(
-						df
-							.filter((FilterFunction<String>) s -> s.startsWith(clazz.getName()))
-							.map(
-								(MapFunction<String, String>) s -> StringUtils.substringAfter(s, "|"),
-								Encoders.STRING()));
-
-				if (filterInvisible) {
-					entityDF = entityDF.filter("dataInfo.invisible != true");
-				}
-
-				entityDF
-					.write()
-					.mode(SaveMode.Overwrite)
-					.option("compression", "gzip")
-					.json(outputPath + "/" + entityType);
-			}
-		});
-	}
-}
--- a/dhp-common/src/main/java/eu/dnetlib/dhp/oa/merge/GroupEntitiesSparkJob.java
+++ b/dhp-common/src/main/java/eu/dnetlib/dhp/oa/merge/GroupEntitiesSparkJob.java
@ -2,36 +2,28 @@
 package eu.dnetlib.dhp.oa.merge;

 import static eu.dnetlib.dhp.common.SparkSessionSupport.runWithSparkSession;
-import static eu.dnetlib.dhp.utils.DHPUtils.toSeq;
+import static org.apache.spark.sql.functions.col;
+import static org.apache.spark.sql.functions.when;

-import java.io.IOException;
-import java.util.List;
-import java.util.Objects;
+import java.util.Map;
 import java.util.Optional;
+import java.util.concurrent.ExecutionException;
+import java.util.concurrent.ForkJoinPool;
 import java.util.stream.Collectors;

 import org.apache.commons.io.IOUtils;
-import org.apache.commons.lang3.StringUtils;
 import org.apache.spark.SparkConf;
-import org.apache.spark.api.java.JavaSparkContext;
-import org.apache.spark.api.java.function.FilterFunction;
 import org.apache.spark.api.java.function.MapFunction;
+import org.apache.spark.api.java.function.ReduceFunction;
 import org.apache.spark.sql.*;
-import org.apache.spark.sql.expressions.Aggregator;
 import org.slf4j.Logger;
 import org.slf4j.LoggerFactory;

-import com.fasterxml.jackson.databind.DeserializationFeature;
-import com.fasterxml.jackson.databind.ObjectMapper;
-import com.jayway.jsonpath.Configuration;
-import com.jayway.jsonpath.DocumentContext;
-import com.jayway.jsonpath.JsonPath;
-import com.jayway.jsonpath.Option;
-
 import eu.dnetlib.dhp.application.ArgumentApplicationParser;
 import eu.dnetlib.dhp.common.HdfsSupport;
+import eu.dnetlib.dhp.schema.common.EntityType;
 import eu.dnetlib.dhp.schema.common.ModelSupport;
-import eu.dnetlib.dhp.schema.oaf.*;
+import eu.dnetlib.dhp.schema.oaf.OafEntity;
 import eu.dnetlib.dhp.schema.oaf.utils.OafMapperUtils;
 import scala.Tuple2;

@ -39,13 +31,9 @@ import scala.Tuple2;
 * Groups the graph content by entity identifier to ensure ID uniqueness
 */
 public class GroupEntitiesSparkJob {
-
 	private static final Logger log = LoggerFactory.getLogger(GroupEntitiesSparkJob.class);

-	private static final String ID_JPATH = "$.id";
-
-	private static final ObjectMapper OBJECT_MAPPER = new ObjectMapper()
-		.configure(DeserializationFeature.FAIL_ON_UNKNOWN_PROPERTIES, false);
+	private static final Encoder<OafEntity> OAFENTITY_KRYO_ENC = Encoders.kryo(OafEntity.class);

 	public static void main(String[] args) throws Exception {

@ -66,9 +54,15 @@ public class GroupEntitiesSparkJob {
 		String graphInputPath = parser.get("graphInputPath");
 		log.info("graphInputPath: {}", graphInputPath);

+		String checkpointPath = parser.get("checkpointPath");
+		log.info("checkpointPath: {}", checkpointPath);
+
 		String outputPath = parser.get("outputPath");
 		log.info("outputPath: {}", outputPath);

+		boolean filterInvisible = Boolean.valueOf(parser.get("filterInvisible"));
+		log.info("filterInvisible: {}", filterInvisible);
+
 		SparkConf conf = new SparkConf();
 		conf.set("spark.serializer", "org.apache.spark.serializer.KryoSerializer");
 		conf.registerKryoClasses(ModelSupport.getOafModelClasses());
@ -77,127 +71,96 @@ public class GroupEntitiesSparkJob {
 			conf,
 			isSparkSessionManaged,
 			spark -> {
-				HdfsSupport.remove(outputPath, spark.sparkContext().hadoopConfiguration());
-				groupEntities(spark, graphInputPath, outputPath);
+				HdfsSupport.remove(checkpointPath, spark.sparkContext().hadoopConfiguration());
+				groupEntities(spark, graphInputPath, checkpointPath, outputPath, filterInvisible);
 			});
 	}

 	private static void groupEntities(
 		SparkSession spark,
 		String inputPath,
-		String outputPath) {
+		String checkpointPath,
+		String outputPath,
+		boolean filterInvisible) {

-		final TypedColumn<OafEntity, OafEntity> aggregator = new GroupingAggregator().toColumn();
-		final JavaSparkContext sc = JavaSparkContext.fromSparkContext(spark.sparkContext());
-		spark
-			.read()
-			.textFile(toSeq(listEntityPaths(inputPath, sc)))
-			.map((MapFunction<String, OafEntity>) GroupEntitiesSparkJob::parseOaf, Encoders.kryo(OafEntity.class))
-			.filter((FilterFunction<OafEntity>) e -> StringUtils.isNotBlank(ModelSupport.idFn().apply(e)))
-			.groupByKey((MapFunction<OafEntity, String>) oaf -> ModelSupport.idFn().apply(oaf), Encoders.STRING())
-			.agg(aggregator)
+		Dataset<OafEntity> allEntities = spark.emptyDataset(OAFENTITY_KRYO_ENC);
+
+		for (Map.Entry<EntityType, Class> e : ModelSupport.entityTypes.entrySet()) {
+			String entity = e.getKey().name();
+			Class<? extends OafEntity> entityClass = e.getValue();
+			String entityInputPath = inputPath + "/" + entity;
+
+			if (!HdfsSupport.exists(entityInputPath, spark.sparkContext().hadoopConfiguration())) {
+				continue;
+			}
+
+			allEntities = allEntities
+				.union(
+					((Dataset<OafEntity>) spark
+						.read()
+						.schema(Encoders.bean(entityClass).schema())
+						.json(entityInputPath)
+						.filter("length(id) > 0")
+						.as(Encoders.bean(entityClass)))
+							.map((MapFunction<OafEntity, OafEntity>) r -> r, OAFENTITY_KRYO_ENC));
+		}
+
+		Dataset<?> groupedEntities = allEntities
+			.groupByKey((MapFunction<OafEntity, String>) OafEntity::getId, Encoders.STRING())
+			.reduceGroups((ReduceFunction<OafEntity>) (b, a) -> OafMapperUtils.mergeEntities(b, a))
 			.map(
-				(MapFunction<Tuple2<String, OafEntity>, String>) t -> t._2().getClass().getName() +
-					"|" + OBJECT_MAPPER.writeValueAsString(t._2()),
-				Encoders.STRING())
+				(MapFunction<Tuple2<String, OafEntity>, Tuple2<String, OafEntity>>) t -> new Tuple2(
+					t._2().getClass().getName(), t._2()),
+				Encoders.tuple(Encoders.STRING(), OAFENTITY_KRYO_ENC));
+
+		// pivot on "_1" (classname of the entity)
+		// created columns containing only entities of the same class
+		for (Map.Entry<EntityType, Class> e : ModelSupport.entityTypes.entrySet()) {
+			String entity = e.getKey().name();
+			Class<? extends OafEntity> entityClass = e.getValue();
+
+			groupedEntities = groupedEntities
+				.withColumn(
+					entity,
+					when(col("_1").equalTo(entityClass.getName()), col("_2")));
+		}
+
+		groupedEntities
+			.drop("_1", "_2")
 			.write()
-			.option("compression", "gzip")
 			.mode(SaveMode.Overwrite)
-			.text(outputPath);
-	}
+			.option("compression", "gzip")
+			.save(checkpointPath);

-	public static class GroupingAggregator extends Aggregator<OafEntity, OafEntity, OafEntity> {
+		ForkJoinPool parPool = new ForkJoinPool(ModelSupport.entityTypes.size());

-		@Override
-		public OafEntity zero() {
-			return null;
-		}
-
-		@Override
-		public OafEntity reduce(OafEntity b, OafEntity a) {
-			return mergeAndGet(b, a);
-		}
-
-		private OafEntity mergeAndGet(OafEntity b, OafEntity a) {
-			if (Objects.nonNull(a) && Objects.nonNull(b)) {
-				return OafMapperUtils.mergeEntities(b, a);
-			}
-			return Objects.isNull(a) ? b : a;
-		}
-
-		@Override
-		public OafEntity merge(OafEntity b, OafEntity a) {
-			return mergeAndGet(b, a);
-		}
-
-		@Override
-		public OafEntity finish(OafEntity j) {
-			return j;
-		}
-
-		@Override
-		public Encoder<OafEntity> bufferEncoder() {
-			return Encoders.kryo(OafEntity.class);
-		}
-
-		@Override
-		public Encoder<OafEntity> outputEncoder() {
-			return Encoders.kryo(OafEntity.class);
-		}
-
-	}
-
-	private static OafEntity parseOaf(String s) {
-
-		DocumentContext dc = JsonPath
-			.parse(s, Configuration.defaultConfiguration().addOptions(Option.SUPPRESS_EXCEPTIONS));
-		final String id = dc.read(ID_JPATH);
-		if (StringUtils.isNotBlank(id)) {
-
-			String prefix = StringUtils.substringBefore(id, "|");
-			switch (prefix) {
-				case "10":
-					return parse(s, Datasource.class);
-				case "20":
-					return parse(s, Organization.class);
-				case "40":
-					return parse(s, Project.class);
-				case "50":
-					String resultType = dc.read("$.resulttype.classid");
-					switch (resultType) {
-						case "publication":
-							return parse(s, Publication.class);
-						case "dataset":
-							return parse(s, eu.dnetlib.dhp.schema.oaf.Dataset.class);
-						case "software":
-							return parse(s, Software.class);
-						case "other":
-							return parse(s, OtherResearchProduct.class);
-						default:
-							throw new IllegalArgumentException(String.format("invalid resultType: '%s'", resultType));
-					}
-				default:
-					throw new IllegalArgumentException(String.format("invalid id prefix: '%s'", prefix));
-			}
-		} else {
-			throw new IllegalArgumentException(String.format("invalid oaf: '%s'", s));
-		}
-	}
-
-	private static <T extends OafEntity> OafEntity parse(String s, Class<T> clazz) {
-		try {
-			return OBJECT_MAPPER.readValue(s, clazz);
-		} catch (IOException e) {
-			throw new IllegalArgumentException(e);
-		}
-	}
-
-	private static List<String> listEntityPaths(String inputPath, JavaSparkContext sc) {
-		return HdfsSupport
-			.listFiles(inputPath, sc.hadoopConfiguration())
+		ModelSupport.entityTypes
+			.entrySet()
 			.stream()
-			.filter(f -> !f.toLowerCase().contains("relation"))
-			.collect(Collectors.toList());
-	}
+			.map(e -> parPool.submit(() -> {
+				String entity = e.getKey().name();
+				Class<? extends OafEntity> entityClass = e.getValue();

+				spark
+					.read()
+					.load(checkpointPath)
+					.select(col(entity).as("value"))
+					.filter("value IS NOT NULL")
+					.as(OAFENTITY_KRYO_ENC)
+					.map((MapFunction<OafEntity, OafEntity>) r -> r, (Encoder<OafEntity>) Encoders.bean(entityClass))
+					.filter(filterInvisible ? "dataInfo.invisible != TRUE" : "TRUE")
+					.write()
+					.mode(SaveMode.Overwrite)
+					.option("compression", "gzip")
+					.json(outputPath + "/" + entity);
+			}))
+			.collect(Collectors.toList())
+			.forEach(t -> {
+				try {
+					t.get();
+				} catch (InterruptedException | ExecutionException e) {
+					throw new RuntimeException(e);
+				}
+			});
+	}
 }
--- a/dhp-common/src/main/java/eu/dnetlib/dhp/schema/oaf/utils/GraphCleaningFunctions.java
+++ b/dhp-common/src/main/java/eu/dnetlib/dhp/schema/oaf/utils/GraphCleaningFunctions.java
@ -36,6 +36,19 @@ public class GraphCleaningFunctions extends CleaningFunctions {

 	public static final int TITLE_FILTER_RESIDUAL_LENGTH = 5;
 	private static final String NAME_CLEANING_REGEX = "[\\r\\n\\t\\s]+";
+	private static final HashSet<String> PEER_REVIEWED_TYPES = new HashSet<>();
+
+	static {
+		PEER_REVIEWED_TYPES.add("Article");
+		PEER_REVIEWED_TYPES.add("Part of book or chapter of book");
+		PEER_REVIEWED_TYPES.add("Book");
+		PEER_REVIEWED_TYPES.add("Doctoral thesis");
+		PEER_REVIEWED_TYPES.add("Master thesis");
+		PEER_REVIEWED_TYPES.add("Data Paper");
+		PEER_REVIEWED_TYPES.add("Thesis");
+		PEER_REVIEWED_TYPES.add("Bachelor thesis");
+		PEER_REVIEWED_TYPES.add("Conference object");
+	}

 	public static <T extends Oaf> T cleanContext(T value, String contextId, String verifyParam) {
 		if (ModelSupport.isSubClass(value, Result.class)) {
@ -493,6 +506,35 @@ public class GraphCleaningFunctions extends CleaningFunctions {
 						if (Objects.isNull(i.getRefereed()) || StringUtils.isBlank(i.getRefereed().getClassid())) {
 							i.setRefereed(qualifier("0000", "Unknown", ModelConstants.DNET_REVIEW_LEVELS));
 						}
+
+						// from the script from Dimitris
+						if ("0000".equals(i.getRefereed().getClassid())) {
+							final boolean isFromCrossref = Optional
+								.ofNullable(i.getCollectedfrom())
+								.map(KeyValue::getKey)
+								.map(id -> id.equals(ModelConstants.CROSSREF_ID))
+								.orElse(false);
+							final boolean hasDoi = Optional
+								.ofNullable(i.getPid())
+								.map(
+									pid -> pid
+										.stream()
+										.anyMatch(
+											p -> PidType.doi.toString().equals(p.getQualifier().getClassid())))
+								.orElse(false);
+							final boolean isPeerReviewedType = PEER_REVIEWED_TYPES
+								.contains(i.getInstancetype().getClassname());
+							final boolean noOtherLitType = r
+								.getInstance()
+								.stream()
+								.noneMatch(ii -> "Other literature type".equals(ii.getInstancetype().getClassname()));
+							if (isFromCrossref && hasDoi && isPeerReviewedType && noOtherLitType) {
+								i.setRefereed(qualifier("0001", "peerReviewed", ModelConstants.DNET_REVIEW_LEVELS));
+							} else {
+								i.setRefereed(qualifier("0002", "nonPeerReviewed", ModelConstants.DNET_REVIEW_LEVELS));
+							}
+						}
+
 						if (Objects.nonNull(i.getDateofacceptance())) {
 							Optional<String> date = cleanDateField(i.getDateofacceptance());
 							if (date.isPresent()) {
--- a/dhp-common/src/main/java/eu/dnetlib/dhp/schema/oaf/utils/PmidCleaningRule.java
+++ b/dhp-common/src/main/java/eu/dnetlib/dhp/schema/oaf/utils/PmidCleaningRule.java
@ -7,7 +7,7 @@ import java.util.regex.Pattern;
 // https://researchguides.stevens.edu/c.php?g=442331&p=6577176
 public class PmidCleaningRule {

-	public static final Pattern PATTERN = Pattern.compile("[1-9]{1,8}");
+	public static final Pattern PATTERN = Pattern.compile("0*(\\d{1,8})");

 	public static String clean(String pmid) {
 		String s = pmid
@ -17,7 +17,7 @@ public class PmidCleaningRule {
 		final Matcher m = PATTERN.matcher(s);

 		if (m.find()) {
-			return m.group();
+			return m.group(1);
 		}
 		return "";
 	}
--- a/dhp-common/src/main/resources/eu/dnetlib/dhp/oa/merge/dispatch_entities_parameters.json
+++ b/dhp-common/src/main/resources/eu/dnetlib/dhp/oa/merge/dispatch_entities_parameters.json
@ -1,26 +0,0 @@
-[
-  {
-    "paramName": "issm",
-    "paramLongName": "isSparkSessionManaged",
-    "paramDescription": "when true will stop SparkSession after job execution",
-    "paramRequired": false
-  },
-  {
-    "paramName": "i",
-    "paramLongName": "inputPath",
-    "paramDescription": "the source path",
-    "paramRequired": true
-  },
-  {
-    "paramName": "o",
-    "paramLongName": "outputPath",
-    "paramDescription": "path of the output graph",
-    "paramRequired": true
-  },
-  {
-    "paramName": "fi",
-    "paramLongName": "filterInvisible",
-    "paramDescription": "if true filters out invisible entities",
-    "paramRequired": true
-  }
-]
--- a/dhp-common/src/main/resources/eu/dnetlib/dhp/oa/merge/group_graph_entities_parameters.json
+++ b/dhp-common/src/main/resources/eu/dnetlib/dhp/oa/merge/group_graph_entities_parameters.json
@ -8,13 +8,25 @@
  {
    "paramName": "gin",
    "paramLongName": "graphInputPath",
-    "paramDescription": "the graph root path",
+    "paramDescription": "the input graph root path",
+    "paramRequired": true
+  },
+  {
+    "paramName": "cp",
+    "paramLongName": "checkpointPath",
+    "paramDescription": "checkpoint directory",
    "paramRequired": true
  },
  {
    "paramName": "out",
    "paramLongName": "outputPath",
-    "paramDescription": "the output merged graph root path",
+    "paramDescription": "the output graph root path",
+    "paramRequired": true
+  },
+  {
+    "paramName": "fi",
+    "paramLongName": "filterInvisible",
+    "paramDescription": "if true filters out invisible entities",
    "paramRequired": true
  }
 ]
--- a/dhp-common/src/test/java/eu/dnetlib/dhp/schema/oaf/utils/PmidCleaningRuleTest.java
+++ b/dhp-common/src/test/java/eu/dnetlib/dhp/schema/oaf/utils/PmidCleaningRuleTest.java
@ -9,10 +9,16 @@ class PmidCleaningRuleTest {

 	@Test
 	void testCleaning() {
+		// leading zeros are removed
 		assertEquals("1234", PmidCleaningRule.clean("01234"));
+		// tolerant to spaces in the middle
 		assertEquals("1234567", PmidCleaningRule.clean("0123 4567"));
+		// stop parsing at first not numerical char
 		assertEquals("123", PmidCleaningRule.clean("0123x4567"));
+		// invalid id leading to empty result
 		assertEquals("", PmidCleaningRule.clean("abc"));
+		// valid id with zeroes in the number
+		assertEquals("20794075", PmidCleaningRule.clean("20794075"));
 	}

 }
--- a/dhp-pace-core/src/main/java/eu/dnetlib/pace/model/SparkModel.scala
+++ b/dhp-pace-core/src/main/java/eu/dnetlib/pace/model/SparkModel.scala
@ -81,7 +81,7 @@ case class SparkModel(conf: DedupConfig) {
              MapDocumentUtil.truncateList(
                MapDocumentUtil.getJPathList(fdef.getPath, documentContext, fdef.getType),
                fdef.getSize
-              ).toArray
+              ).asScala

            case Type.StringConcat =>
              val jpaths = CONCAT_REGEX.split(fdef.getPath)
--- a/dhp-pace-core/src/main/java/eu/dnetlib/pace/util/DiffPatchMatch.java
+++ b/dhp-pace-core/src/main/java/eu/dnetlib/pace/util/DiffPatchMatch.java
@ -1,6 +1,23 @@

 package eu.dnetlib.pace.util;

+/*
+ * Diff Match and Patch
+ * Copyright 2018 The diff-match-patch Authors.
+ * https://github.com/google/diff-match-patch
+ *
+ * Licensed under the Apache License, Version 2.0 (the "License");
+ * you may not use this file except in compliance with the License.
+ * You may obtain a copy of the License at
+ *
+ *   http://www.apache.org/licenses/LICENSE-2.0
+ *
+ * Unless required by applicable law or agreed to in writing, software
+ * distributed under the License is distributed on an "AS IS" BASIS,
+ * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
+ * See the License for the specific language governing permissions and
+ * limitations under the License.
+ */
 /*
 * Diff Match and Patch
 * Copyright 2018 The diff-match-patch Authors.
--- a/dhp-pace-core/src/main/java/eu/dnetlib/pace/util/MapDocumentUtil.java
+++ b/dhp-pace-core/src/main/java/eu/dnetlib/pace/util/MapDocumentUtil.java
@ -117,6 +117,11 @@ public class MapDocumentUtil {
 			return result;
 		}

+		if (type == Type.List && jresult instanceof List) {
+			((List<?>) jresult).forEach(x -> result.add(x.toString()));
+			return result;
+		}
+
 		if (jresult instanceof JSONArray) {
 			((JSONArray) jresult).forEach(it -> {
 				try {
--- a/dhp-workflows/dhp-actionmanager/README.md
+++ b/dhp-workflows/dhp-actionmanager/README.md
@ -0,0 +1,72 @@
+# Action Management Framework
+
+This module implements the oozie workflow for the integration of pre-built contents into the OpenAIRE Graph.
+
+Such contents can be 
+
+* brand new, non-existing records to be introduced as nodes of the graph
+* updates (or enrichment) for records that does exist in the graph (e.g. a new subject term for a publication)
+* relations among existing nodes
+
+The actionset contents are organised into logical containers, each of them can contain multiple versions contents and is characterised by
+
+* a name
+* an identifier
+* the paths on HDFS where each version of the contents is stored
+
+Each version is then characterised by 
+
+* the creation date
+* the last update date
+* the indication where it is the latest one or it is an expired version, candidate for garbage collection
+
+## ActionSet serialization
+
+Each actionset version contains records compliant to the graph internal data model, i.e. subclasses of `eu.dnetlib.dhp.schema.oaf.Oaf`,
+defined in the external schemas module
+
+```
+<dependency>
+    <groupId>eu.dnetlib.dhp</groupId>
+    <artifactId>${dhp-schemas.artifact}</artifactId>
+    <version>${dhp-schemas.version}</version>
+</dependency>
+```
+
+When the actionset contains a relationship, the model class to use is `eu.dnetlib.dhp.schema.oaf.Relation`, otherwise 
+when the actionset contains an entity, it is a `eu.dnetlib.dhp.schema.oaf.OafEntity` or one of its subclasses 
+`Datasource`, `Organization`, `Project`, `Result` (or one of its subclasses `Publication`, `Dataset`, etc...). 
+
+Then, each OpenAIRE Graph model class instance must be wrapped using the class `eu.dnetlib.dhp.schema.action.AtomicAction`, a generic 
+container that defines two attributes
+
+* `T payload` the OpenAIRE Graph class instance containing the data;
+* `Class<T> clazz` must contain the class whose instance is contained in the payload.
+
+Each AtomicAction can be then serialised in JSON format using `com.fasterxml.jackson.databind.ObjectMapper` from
+
+```
+<dependency>
+    <groupId>com.fasterxml.jackson.core</groupId>
+    <artifactId>jackson-databind</artifactId>
+    <version>${dhp.jackson.version}</version>
+</dependency>
+```
+
+Then, the JSON serialization must be stored as a GZip compressed sequence file (`org.apache.hadoop.mapred.SequenceFileOutputFormat`). 
+As such, it contains a set of tuples, a key and a value defined as `org.apache.hadoop.io.Text` where
+
+* the `key` must be set to the class canonical name contained in the `AtomicAction`;
+* the `value` must be set to the AtomicAction JSON serialization.
+
+The following snippet provides an example of how create an actionset version of Relation records:
+
+```
+  rels // JavaRDD<Relation>
+    .map(relation -> new AtomicAction<Relation>(Relation.class, relation))
+    .mapToPair(
+        aa -> new Tuple2<>(new Text(aa.getClazz().getCanonicalName()),
+            new Text(OBJECT_MAPPER.writeValueAsString(aa))))
+    .saveAsHadoopFile(outputPath, Text.class, Text.class, SequenceFileOutputFormat.class, GzipCodec.class);
+```
+
--- a/dhp-workflows/dhp-aggregation/src/main/java/eu/dnetlib/dhp/actionmanager/Constants.java
+++ b/dhp-workflows/dhp-aggregation/src/main/java/eu/dnetlib/dhp/actionmanager/Constants.java
@ -40,6 +40,7 @@ public class Constants {
 	public static final String SDG_CLASS_NAME = "Sustainable Development Goals";

 	public static final String NULL = "NULL";
+	public static final String NA = "N/A";

 	public static final ObjectMapper OBJECT_MAPPER = new ObjectMapper();

@ -61,10 +62,16 @@ public class Constants {
 			.map((MapFunction<String, R>) value -> OBJECT_MAPPER.readValue(value, clazz), Encoders.bean(clazz));
 	}

-	public static Subject getSubject(String sbj, String classid, String classname,
-		String diqualifierclassid) {
-		if (sbj == null || sbj.equals(NULL))
+	public static Subject getSubject(String sbj, String classid, String classname, String diqualifierclassid,
+		Boolean split) {
+		if (sbj == null || sbj.equals(NULL) || sbj.startsWith(NA))
 			return null;
+		String trust = "";
+		String subject = sbj;
+		if (split) {
+			sbj = subject.split("@@")[0];
+			trust = subject.split("@@")[1];
+		}
 		Subject s = new Subject();
 		s.setValue(sbj);
 		s
@ -89,9 +96,14 @@ public class Constants {
 								UPDATE_CLASS_NAME,
 								ModelConstants.DNET_PROVENANCE_ACTIONS,
 								ModelConstants.DNET_PROVENANCE_ACTIONS),
-						""));
+						trust));

 		return s;
+	}
+
+	public static Subject getSubject(String sbj, String classid, String classname,
+		String diqualifierclassid) {
+		return getSubject(sbj, classid, classname, diqualifierclassid, false);

 	}

--- a/dhp-workflows/dhp-aggregation/src/main/java/eu/dnetlib/dhp/actionmanager/bipfinder/SparkAtomicActionScoreJob.java
+++ b/dhp-workflows/dhp-aggregation/src/main/java/eu/dnetlib/dhp/actionmanager/bipfinder/SparkAtomicActionScoreJob.java
@ -95,7 +95,7 @@ public class SparkAtomicActionScoreJob implements Serializable {

 		return projectScores.map((MapFunction<BipProjectModel, Project>) bipProjectScores -> {
 			Project project = new Project();
-			project.setId(bipProjectScores.getProjectId());
+			// project.setId(bipProjectScores.getProjectId());
 			project.setMeasures(bipProjectScores.toMeasures());
 			return project;
 		}, Encoders.bean(Project.class))
--- a/dhp-workflows/dhp-aggregation/src/main/java/eu/dnetlib/dhp/actionmanager/createunresolvedentities/GetFOSSparkJob.java
+++ b/dhp-workflows/dhp-aggregation/src/main/java/eu/dnetlib/dhp/actionmanager/createunresolvedentities/GetFOSSparkJob.java
@ -75,9 +75,12 @@ public class GetFOSSparkJob implements Serializable {
 		fosData.map((MapFunction<Row, FOSDataModel>) r -> {
 			FOSDataModel fosDataModel = new FOSDataModel();
 			fosDataModel.setDoi(r.getString(0).toLowerCase());
-			fosDataModel.setLevel1(r.getString(1));
-			fosDataModel.setLevel2(r.getString(2));
-			fosDataModel.setLevel3(r.getString(3));
+			fosDataModel.setLevel1(r.getString(2));
+			fosDataModel.setLevel2(r.getString(3));
+			fosDataModel.setLevel3(r.getString(4));
+			fosDataModel.setLevel4(r.getString(5));
+			fosDataModel.setScoreL3(String.valueOf(r.getDouble(6)));
+			fosDataModel.setScoreL4(String.valueOf(r.getDouble(7)));
 			return fosDataModel;
 		}, Encoders.bean(FOSDataModel.class))
 			.write()
--- a/dhp-workflows/dhp-aggregation/src/main/java/eu/dnetlib/dhp/actionmanager/createunresolvedentities/PrepareBipFinder.java
+++ b/dhp-workflows/dhp-aggregation/src/main/java/eu/dnetlib/dhp/actionmanager/createunresolvedentities/PrepareBipFinder.java
@ -1,178 +0,0 @@
-
-package eu.dnetlib.dhp.actionmanager.createunresolvedentities;
-
-import static eu.dnetlib.dhp.actionmanager.Constants.*;
-import static eu.dnetlib.dhp.actionmanager.Constants.UPDATE_CLASS_NAME;
-import static eu.dnetlib.dhp.common.SparkSessionSupport.runWithSparkSession;
-
-import java.io.Serializable;
-import java.util.Arrays;
-import java.util.List;
-import java.util.Optional;
-import java.util.stream.Collectors;
-
-import org.apache.commons.io.IOUtils;
-import org.apache.spark.SparkConf;
-import org.apache.spark.api.java.JavaRDD;
-import org.apache.spark.api.java.JavaSparkContext;
-import org.apache.spark.api.java.function.MapFunction;
-import org.apache.spark.sql.Encoders;
-import org.apache.spark.sql.SaveMode;
-import org.apache.spark.sql.SparkSession;
-import org.slf4j.Logger;
-import org.slf4j.LoggerFactory;
-
-import com.fasterxml.jackson.databind.ObjectMapper;
-
-import eu.dnetlib.dhp.actionmanager.bipmodel.BipScore;
-import eu.dnetlib.dhp.actionmanager.bipmodel.score.deserializers.BipResultModel;
-import eu.dnetlib.dhp.application.ArgumentApplicationParser;
-import eu.dnetlib.dhp.common.HdfsSupport;
-import eu.dnetlib.dhp.schema.common.ModelConstants;
-import eu.dnetlib.dhp.schema.oaf.Instance;
-import eu.dnetlib.dhp.schema.oaf.KeyValue;
-import eu.dnetlib.dhp.schema.oaf.Measure;
-import eu.dnetlib.dhp.schema.oaf.Result;
-import eu.dnetlib.dhp.schema.oaf.utils.CleaningFunctions;
-import eu.dnetlib.dhp.schema.oaf.utils.OafMapperUtils;
-import eu.dnetlib.dhp.utils.DHPUtils;
-
-public class PrepareBipFinder implements Serializable {
-
-	private static final Logger log = LoggerFactory.getLogger(PrepareBipFinder.class);
-	private static final ObjectMapper OBJECT_MAPPER = new ObjectMapper();
-
-	public static void main(String[] args) throws Exception {
-
-		String jsonConfiguration = IOUtils
-			.toString(
-				PrepareBipFinder.class
-					.getResourceAsStream(
-						"/eu/dnetlib/dhp/actionmanager/createunresolvedentities/prepare_parameters.json"));
-
-		final ArgumentApplicationParser parser = new ArgumentApplicationParser(jsonConfiguration);
-
-		parser.parseArgument(args);
-
-		Boolean isSparkSessionManaged = Optional
-			.ofNullable(parser.get("isSparkSessionManaged"))
-			.map(Boolean::valueOf)
-			.orElse(Boolean.TRUE);
-
-		log.info("isSparkSessionManaged: {}", isSparkSessionManaged);
-
-		final String sourcePath = parser.get("sourcePath");
-		log.info("sourcePath {}: ", sourcePath);
-
-		final String outputPath = parser.get("outputPath");
-		log.info("outputPath {}: ", outputPath);
-
-		SparkConf conf = new SparkConf();
-
-		runWithSparkSession(
-			conf,
-			isSparkSessionManaged,
-			spark -> {
-				HdfsSupport.remove(outputPath, spark.sparkContext().hadoopConfiguration());
-				prepareResults(spark, sourcePath, outputPath);
-			});
-	}
-
-	private static void prepareResults(SparkSession spark, String inputPath, String outputPath) {
-
-		final JavaSparkContext sc = JavaSparkContext.fromSparkContext(spark.sparkContext());
-
-		JavaRDD<BipResultModel> bipDeserializeJavaRDD = sc
-			.textFile(inputPath)
-			.map(item -> OBJECT_MAPPER.readValue(item, BipResultModel.class));
-
-		spark
-			.createDataset(bipDeserializeJavaRDD.flatMap(entry -> entry.keySet().stream().map(key -> {
-				BipScore bs = new BipScore();
-				bs.setId(key);
-				bs.setScoreList(entry.get(key));
-
-				return bs;
-			}).collect(Collectors.toList()).iterator()).rdd(), Encoders.bean(BipScore.class))
-			.map((MapFunction<BipScore, Result>) v -> {
-				Result r = new Result();
-				final String cleanedPid = CleaningFunctions.normalizePidValue(DOI, v.getId());
-
-				r.setId(DHPUtils.generateUnresolvedIdentifier(v.getId(), DOI));
-				Instance inst = new Instance();
-				inst.setMeasures(getMeasure(v));
-
-				inst
-					.setPid(
-						Arrays
-							.asList(
-								OafMapperUtils
-									.structuredProperty(
-										cleanedPid,
-										OafMapperUtils
-											.qualifier(
-												DOI, DOI_CLASSNAME,
-												ModelConstants.DNET_PID_TYPES,
-												ModelConstants.DNET_PID_TYPES),
-										null)));
-				r.setInstance(Arrays.asList(inst));
-				r
-					.setDataInfo(
-						OafMapperUtils
-							.dataInfo(
-								false, null, true,
-								false,
-								OafMapperUtils
-									.qualifier(
-										ModelConstants.PROVENANCE_ENRICH,
-										null,
-										ModelConstants.DNET_PROVENANCE_ACTIONS,
-										ModelConstants.DNET_PROVENANCE_ACTIONS),
-								null));
-				return r;
-			}, Encoders.bean(Result.class))
-			.write()
-			.mode(SaveMode.Overwrite)
-			.option("compression", "gzip")
-			.json(outputPath + "/bip");
-	}
-
-	private static List<Measure> getMeasure(BipScore value) {
-		return value
-			.getScoreList()
-			.stream()
-			.map(score -> {
-				Measure m = new Measure();
-				m.setId(score.getId());
-				m
-					.setUnit(
-						score
-							.getUnit()
-							.stream()
-							.map(unit -> {
-								KeyValue kv = new KeyValue();
-								kv.setValue(unit.getValue());
-								kv.setKey(unit.getKey());
-								kv
-									.setDataInfo(
-										OafMapperUtils
-											.dataInfo(
-												false,
-												UPDATE_DATA_INFO_TYPE,
-												true,
-												false,
-												OafMapperUtils
-													.qualifier(
-														UPDATE_MEASURE_BIP_CLASS_ID,
-														UPDATE_CLASS_NAME,
-														ModelConstants.DNET_PROVENANCE_ACTIONS,
-														ModelConstants.DNET_PROVENANCE_ACTIONS),
-												""));
-								return kv;
-							})
-							.collect(Collectors.toList()));
-				return m;
-			})
-			.collect(Collectors.toList());
-	}
-}
--- a/dhp-workflows/dhp-aggregation/src/main/java/eu/dnetlib/dhp/actionmanager/createunresolvedentities/PrepareFOSSparkJob.java
+++ b/dhp-workflows/dhp-aggregation/src/main/java/eu/dnetlib/dhp/actionmanager/createunresolvedentities/PrepareFOSSparkJob.java
@ -78,12 +78,20 @@ public class PrepareFOSSparkJob implements Serializable {
 				HashSet<String> level1 = new HashSet<>();
 				HashSet<String> level2 = new HashSet<>();
 				HashSet<String> level3 = new HashSet<>();
-				addLevels(level1, level2, level3, first);
-				it.forEachRemaining(v -> addLevels(level1, level2, level3, v));
+				HashSet<String> level4 = new HashSet<>();
+				addLevels(level1, level2, level3, level4, first);
+				it.forEachRemaining(v -> addLevels(level1, level2, level3, level4, v));
 				List<Subject> sbjs = new ArrayList<>();
-				level1.forEach(l -> sbjs.add(getSubject(l, FOS_CLASS_ID, FOS_CLASS_NAME, UPDATE_SUBJECT_FOS_CLASS_ID)));
-				level2.forEach(l -> sbjs.add(getSubject(l, FOS_CLASS_ID, FOS_CLASS_NAME, UPDATE_SUBJECT_FOS_CLASS_ID)));
-				level3.forEach(l -> sbjs.add(getSubject(l, FOS_CLASS_ID, FOS_CLASS_NAME, UPDATE_SUBJECT_FOS_CLASS_ID)));
+				level1
+					.forEach(l -> add(sbjs, getSubject(l, FOS_CLASS_ID, FOS_CLASS_NAME, UPDATE_SUBJECT_FOS_CLASS_ID)));
+				level2
+					.forEach(l -> add(sbjs, getSubject(l, FOS_CLASS_ID, FOS_CLASS_NAME, UPDATE_SUBJECT_FOS_CLASS_ID)));
+				level3
+					.forEach(
+						l -> add(sbjs, getSubject(l, FOS_CLASS_ID, FOS_CLASS_NAME, UPDATE_SUBJECT_FOS_CLASS_ID, true)));
+				level4
+					.forEach(
+						l -> add(sbjs, getSubject(l, FOS_CLASS_ID, FOS_CLASS_NAME, UPDATE_SUBJECT_FOS_CLASS_ID, true)));
 				r.setSubject(sbjs);
 				r
 					.setDataInfo(
@ -106,11 +114,18 @@ public class PrepareFOSSparkJob implements Serializable {
 			.json(outputPath + "/fos");
 	}

+	private static void add(List<Subject> sbsjs, Subject sbj) {
+		if (sbj != null)
+			sbsjs.add(sbj);
+	}
+
 	private static void addLevels(HashSet<String> level1, HashSet<String> level2, HashSet<String> level3,
+		HashSet<String> level4,
 		FOSDataModel first) {
 		level1.add(first.getLevel1());
 		level2.add(first.getLevel2());
-		level3.add(first.getLevel3());
+		level3.add(first.getLevel3() + "@@" + first.getScoreL3());
+		level4.add(first.getLevel4() + "@@" + first.getScoreL4());
 	}

 }
--- a/dhp-workflows/dhp-aggregation/src/main/java/eu/dnetlib/dhp/actionmanager/createunresolvedentities/SparkSaveUnresolved.java
+++ b/dhp-workflows/dhp-aggregation/src/main/java/eu/dnetlib/dhp/actionmanager/createunresolvedentities/SparkSaveUnresolved.java
@ -69,9 +69,9 @@ public class SparkSaveUnresolved implements Serializable {
 			.mapGroups((MapGroupsFunction<String, Result, Result>) (k, it) -> {
 				Result ret = it.next();
 				it.forEachRemaining(r -> {
-					if (r.getInstance() != null) {
-						ret.setInstance(r.getInstance());
-					}
+//					if (r.getInstance() != null) {
+//						ret.setInstance(r.getInstance());
+//					}
 					if (r.getSubject() != null) {
 						if (ret.getSubject() != null)
 							ret.getSubject().addAll(r.getSubject());
--- a/dhp-workflows/dhp-aggregation/src/main/java/eu/dnetlib/dhp/actionmanager/createunresolvedentities/model/FOSDataModel.java
+++ b/dhp-workflows/dhp-aggregation/src/main/java/eu/dnetlib/dhp/actionmanager/createunresolvedentities/model/FOSDataModel.java
@ -11,21 +11,43 @@ public class FOSDataModel implements Serializable {
 	private String doi;

 	@CsvBindByPosition(position = 1)
+//    @CsvBindByName(column = "doi")
+	private String oaid;
+	@CsvBindByPosition(position = 2)
 //    @CsvBindByName(column = "level1")
 	private String level1;

-	@CsvBindByPosition(position = 2)
+	@CsvBindByPosition(position = 3)
 //    @CsvBindByName(column = "level2")
 	private String level2;

-	@CsvBindByPosition(position = 3)
+	@CsvBindByPosition(position = 4)
 //    @CsvBindByName(column = "level3")
 	private String level3;

+	@CsvBindByPosition(position = 5)
+//    @CsvBindByName(column = "level3")
+	private String level4;
+	@CsvBindByPosition(position = 6)
+	private String scoreL3;
+	@CsvBindByPosition(position = 7)
+	private String scoreL4;
+
 	public FOSDataModel() {

 	}

+	public FOSDataModel(String doi, String level1, String level2, String level3, String level4, String l3score,
+		String l4score) {
+		this.doi = doi;
+		this.level1 = level1;
+		this.level2 = level2;
+		this.level3 = level3;
+		this.level4 = level4;
+		this.scoreL3 = l3score;
+		this.scoreL4 = l4score;
+	}
+
 	public FOSDataModel(String doi, String level1, String level2, String level3) {
 		this.doi = doi;
 		this.level1 = level1;
@ -33,8 +55,41 @@ public class FOSDataModel implements Serializable {
 		this.level3 = level3;
 	}

-	public static FOSDataModel newInstance(String d, String level1, String level2, String level3) {
-		return new FOSDataModel(d, level1, level2, level3);
+	public static FOSDataModel newInstance(String d, String level1, String level2, String level3, String level4,
+		String scorel3, String scorel4) {
+		return new FOSDataModel(d, level1, level2, level3, level4, scorel3, scorel4);
+	}
+
+	public String getOaid() {
+		return oaid;
+	}
+
+	public void setOaid(String oaid) {
+		this.oaid = oaid;
+	}
+
+	public String getLevel4() {
+		return level4;
+	}
+
+	public void setLevel4(String level4) {
+		this.level4 = level4;
+	}
+
+	public String getScoreL3() {
+		return scoreL3;
+	}
+
+	public void setScoreL3(String scoreL3) {
+		this.scoreL3 = scoreL3;
+	}
+
+	public String getScoreL4() {
+		return scoreL4;
+	}
+
+	public void setScoreL4(String scoreL4) {
+		this.scoreL4 = scoreL4;
 	}

 	public String getDoi() {
--- a/dhp-workflows/dhp-aggregation/src/main/java/eu/dnetlib/dhp/actionmanager/opencitations/CreateActionSetSparkJob.java
+++ b/dhp-workflows/dhp-aggregation/src/main/java/eu/dnetlib/dhp/actionmanager/opencitations/CreateActionSetSparkJob.java
@ -10,8 +10,10 @@ import java.util.*;
 import org.apache.commons.cli.ParseException;
 import org.apache.commons.io.IOUtils;
 import org.apache.hadoop.io.Text;
+import org.apache.hadoop.io.compress.GzipCodec;
 import org.apache.hadoop.mapred.SequenceFileOutputFormat;
 import org.apache.spark.SparkConf;
+import org.apache.spark.api.java.JavaPairRDD;
 import org.apache.spark.api.java.function.FilterFunction;
 import org.apache.spark.api.java.function.FlatMapFunction;
 import org.apache.spark.api.java.function.MapFunction;
@ -26,19 +28,29 @@ import eu.dnetlib.dhp.actionmanager.opencitations.model.COCI;
 import eu.dnetlib.dhp.application.ArgumentApplicationParser;
 import eu.dnetlib.dhp.schema.action.AtomicAction;
 import eu.dnetlib.dhp.schema.common.ModelConstants;
-import eu.dnetlib.dhp.schema.common.ModelSupport;
 import eu.dnetlib.dhp.schema.oaf.*;
-import eu.dnetlib.dhp.schema.oaf.utils.CleaningFunctions;
-import eu.dnetlib.dhp.schema.oaf.utils.IdentifierFactory;
+import eu.dnetlib.dhp.schema.oaf.utils.*;
+import eu.dnetlib.dhp.utils.DHPUtils;
 import scala.Tuple2;

 public class CreateActionSetSparkJob implements Serializable {
 	public static final String OPENCITATIONS_CLASSID = "sysimport:crosswalk:opencitations";
 	public static final String OPENCITATIONS_CLASSNAME = "Imported from OpenCitations";
-	private static final String ID_PREFIX = "50|doi_________::";
+
+	// DOI-to-DOI citations
+	public static final String COCI = "COCI";
+
+	// PMID-to-PMID citations
+	public static final String POCI = "POCI";
+
+	private static final String DOI_PREFIX = "50|doi_________::";
+
+	private static final String PMID_PREFIX = "50|pmid________::";
+
 	private static final String TRUST = "0.91";

 	private static final Logger log = LoggerFactory.getLogger(CreateActionSetSparkJob.class);
+
 	private static final ObjectMapper OBJECT_MAPPER = new ObjectMapper();

 	public static void main(final String[] args) throws IOException, ParseException {
@ -62,7 +74,7 @@ public class CreateActionSetSparkJob implements Serializable {
 		log.info("isSparkSessionManaged: {}", isSparkSessionManaged);

 		final String inputPath = parser.get("inputPath");
-		log.info("inputPath {}", inputPath.toString());
+		log.info("inputPath {}", inputPath);

 		final String outputPath = parser.get("outputPath");
 		log.info("outputPath {}", outputPath);
@ -76,41 +88,68 @@ public class CreateActionSetSparkJob implements Serializable {
 		runWithSparkSession(
 			conf,
 			isSparkSessionManaged,
-			spark -> {
-				extractContent(spark, inputPath, outputPath, shouldDuplicateRels);
-			});
+			spark -> extractContent(spark, inputPath, outputPath, shouldDuplicateRels));

 	}

 	private static void extractContent(SparkSession spark, String inputPath, String outputPath,
 		boolean shouldDuplicateRels) {
-		spark
+
+		getTextTextJavaPairRDD(spark, inputPath, shouldDuplicateRels, COCI)
+			.union(getTextTextJavaPairRDD(spark, inputPath, shouldDuplicateRels, POCI))
+			.saveAsHadoopFile(outputPath, Text.class, Text.class, SequenceFileOutputFormat.class, GzipCodec.class);
+	}
+
+	private static JavaPairRDD<Text, Text> getTextTextJavaPairRDD(SparkSession spark, String inputPath,
+		boolean shouldDuplicateRels, String prefix) {
+		return spark
 			.read()
-			.textFile(inputPath + "/*")
+			.textFile(inputPath + "/" + prefix + "/" + prefix + "_JSON/*")
 			.map(
 				(MapFunction<String, COCI>) value -> OBJECT_MAPPER.readValue(value, COCI.class),
 				Encoders.bean(COCI.class))
 			.flatMap(
-				(FlatMapFunction<COCI, Relation>) value -> createRelation(value, shouldDuplicateRels).iterator(),
+				(FlatMapFunction<COCI, Relation>) value -> createRelation(
+					value, shouldDuplicateRels, prefix)
+						.iterator(),
 				Encoders.bean(Relation.class))
-			.filter((FilterFunction<Relation>) value -> value != null)
+			.filter((FilterFunction<Relation>) Objects::nonNull)
 			.toJavaRDD()
 			.map(p -> new AtomicAction(p.getClass(), p))
 			.mapToPair(
 				aa -> new Tuple2<>(new Text(aa.getClazz().getCanonicalName()),
-					new Text(OBJECT_MAPPER.writeValueAsString(aa))))
-			.saveAsHadoopFile(outputPath, Text.class, Text.class, SequenceFileOutputFormat.class);
-
+					new Text(OBJECT_MAPPER.writeValueAsString(aa))));
 	}

-	private static List<Relation> createRelation(COCI value, boolean duplicate) {
+	private static List<Relation> createRelation(COCI value, boolean duplicate, String p) {

 		List<Relation> relationList = new ArrayList<>();
+		String prefix;
+		String citing;
+		String cited;

-		String citing = ID_PREFIX
-			+ IdentifierFactory.md5(CleaningFunctions.normalizePidValue("doi", value.getCiting()));
-		final String cited = ID_PREFIX
-			+ IdentifierFactory.md5(CleaningFunctions.normalizePidValue("doi", value.getCited()));
+		switch (p) {
+			case COCI:
+				prefix = DOI_PREFIX;
+				citing = prefix
+					+ IdentifierFactory
+						.md5(PidCleaner.normalizePidValue(PidType.doi.toString(), value.getCiting()));
+				cited = prefix
+					+ IdentifierFactory
+						.md5(PidCleaner.normalizePidValue(PidType.doi.toString(), value.getCited()));
+				break;
+			case POCI:
+				prefix = PMID_PREFIX;
+				citing = prefix
+					+ IdentifierFactory
+						.md5(PidCleaner.normalizePidValue(PidType.pmid.toString(), value.getCiting()));
+				cited = prefix
+					+ IdentifierFactory
+						.md5(PidCleaner.normalizePidValue(PidType.pmid.toString(), value.getCited()));
+				break;
+			default:
+				throw new IllegalStateException("Invalid prefix: " + p);
+		}

 		if (!citing.equals(cited)) {
 			relationList
@ -120,7 +159,7 @@ public class CreateActionSetSparkJob implements Serializable {
 						cited, ModelConstants.CITES));

 			if (duplicate && value.getCiting().endsWith(".refs")) {
-				citing = ID_PREFIX + IdentifierFactory
+				citing = prefix + IdentifierFactory
 					.md5(
 						CleaningFunctions
 							.normalizePidValue(
@ -132,59 +171,30 @@ public class CreateActionSetSparkJob implements Serializable {
 		return relationList;
 	}

-	private static Collection<Relation> getRelations(String citing, String cited) {
-
-		return Arrays
-			.asList(
-				getRelation(citing, cited, ModelConstants.CITES),
-				getRelation(cited, citing, ModelConstants.IS_CITED_BY));
-	}
-
 	public static Relation getRelation(
 		String source,
 		String target,
-		String relclass) {
-		Relation r = new Relation();
-		r.setCollectedfrom(getCollectedFrom());
-		r.setSource(source);
-		r.setTarget(target);
-		r.setRelClass(relclass);
-		r.setRelType(ModelConstants.RESULT_RESULT);
-		r.setSubRelType(ModelConstants.CITATION);
-		r
-			.setDataInfo(
-				getDataInfo());
-		return r;
-	}
+		String relClass) {

-	public static List<KeyValue> getCollectedFrom() {
-		KeyValue kv = new KeyValue();
-		kv.setKey(ModelConstants.OPENOCITATIONS_ID);
-		kv.setValue(ModelConstants.OPENOCITATIONS_NAME);
-
-		return Arrays.asList(kv);
-	}
-
-	public static DataInfo getDataInfo() {
-		DataInfo di = new DataInfo();
-		di.setInferred(false);
-		di.setDeletedbyinference(false);
-		di.setTrust(TRUST);
-
-		di
-			.setProvenanceaction(
-				getQualifier(OPENCITATIONS_CLASSID, OPENCITATIONS_CLASSNAME, ModelConstants.DNET_PROVENANCE_ACTIONS));
-		return di;
-	}
-
-	public static Qualifier getQualifier(String class_id, String class_name,
-		String qualifierSchema) {
-		Qualifier pa = new Qualifier();
-		pa.setClassid(class_id);
-		pa.setClassname(class_name);
-		pa.setSchemeid(qualifierSchema);
-		pa.setSchemename(qualifierSchema);
-		return pa;
+		return OafMapperUtils
+			.getRelation(
+				source,
+				target,
+				ModelConstants.RESULT_RESULT,
+				ModelConstants.CITATION,
+				relClass,
+				Arrays
+					.asList(
+						OafMapperUtils.keyValue(ModelConstants.OPENOCITATIONS_ID, ModelConstants.OPENOCITATIONS_NAME)),
+				OafMapperUtils
+					.dataInfo(
+						false, null, false, false,
+						OafMapperUtils
+							.qualifier(
+								OPENCITATIONS_CLASSID, OPENCITATIONS_CLASSNAME,
+								ModelConstants.DNET_PROVENANCE_ACTIONS, ModelConstants.DNET_PROVENANCE_ACTIONS),
+						TRUST),
+				null);
 	}

 }
--- a/dhp-workflows/dhp-aggregation/src/main/java/eu/dnetlib/dhp/actionmanager/opencitations/GetOpenCitationsRefs.java
+++ b/dhp-workflows/dhp-aggregation/src/main/java/eu/dnetlib/dhp/actionmanager/opencitations/GetOpenCitationsRefs.java
@ -3,6 +3,7 @@ package eu.dnetlib.dhp.actionmanager.opencitations;

 import java.io.*;
 import java.io.Serializable;
+import java.util.Arrays;
 import java.util.Objects;
 import java.util.zip.GZIPOutputStream;
 import java.util.zip.ZipEntry;
@ -37,7 +38,7 @@ public class GetOpenCitationsRefs implements Serializable {
 		parser.parseArgument(args);

 		final String[] inputFile = parser.get("inputFile").split(";");
-		log.info("inputFile {}", inputFile.toString());
+		log.info("inputFile {}", Arrays.asList(inputFile));

 		final String workingPath = parser.get("workingPath");
 		log.info("workingPath {}", workingPath);
@ -45,6 +46,9 @@ public class GetOpenCitationsRefs implements Serializable {
 		final String hdfsNameNode = parser.get("hdfsNameNode");
 		log.info("hdfsNameNode {}", hdfsNameNode);

+		final String prefix = parser.get("prefix");
+		log.info("prefix {}", prefix);
+
 		Configuration conf = new Configuration();
 		conf.set("fs.defaultFS", hdfsNameNode);

@ -53,30 +57,31 @@ public class GetOpenCitationsRefs implements Serializable {
 		GetOpenCitationsRefs ocr = new GetOpenCitationsRefs();

 		for (String file : inputFile) {
-			ocr.doExtract(workingPath + "/Original/" + file, workingPath, fileSystem);
+			ocr.doExtract(workingPath + "/Original/" + file, workingPath, fileSystem, prefix);
 		}

 	}

-	private void doExtract(String inputFile, String workingPath, FileSystem fileSystem)
+	private void doExtract(String inputFile, String workingPath, FileSystem fileSystem, String prefix)
 		throws IOException {

 		final Path path = new Path(inputFile);

 		FSDataInputStream oc_zip = fileSystem.open(path);

-		int count = 1;
+		// int count = 1;
 		try (ZipInputStream zis = new ZipInputStream(oc_zip)) {
 			ZipEntry entry = null;
 			while ((entry = zis.getNextEntry()) != null) {

 				if (!entry.isDirectory()) {
 					String fileName = entry.getName();
-					fileName = fileName.substring(0, fileName.indexOf("T")) + "_" + count;
-					count++;
+					// fileName = fileName.substring(0, fileName.indexOf("T")) + "_" + count;
+					fileName = fileName.substring(0, fileName.lastIndexOf("."));
+					// count++;
 					try (
 						FSDataOutputStream out = fileSystem
-							.create(new Path(workingPath + "/COCI/" + fileName + ".gz"));
+							.create(new Path(workingPath + "/" + prefix + "/" + fileName + ".gz"));
 						GZIPOutputStream gzipOs = new GZIPOutputStream(new BufferedOutputStream(out))) {

 						IOUtils.copy(zis, gzipOs);
--- a/dhp-workflows/dhp-aggregation/src/main/java/eu/dnetlib/dhp/actionmanager/opencitations/ReadCOCI.java
+++ b/dhp-workflows/dhp-aggregation/src/main/java/eu/dnetlib/dhp/actionmanager/opencitations/ReadCOCI.java
@ -7,6 +7,7 @@ import static eu.dnetlib.dhp.common.SparkSessionSupport.runWithSparkSession;

 import java.io.IOException;
 import java.io.Serializable;
+import java.util.Arrays;
 import java.util.Optional;

 import org.apache.commons.io.IOUtils;
@ -42,13 +43,16 @@ public class ReadCOCI implements Serializable {
 		log.info("outputPath: {}", outputPath);

 		final String[] inputFile = parser.get("inputFile").split(";");
-		log.info("inputFile {}", inputFile.toString());
+		log.info("inputFile {}", Arrays.asList(inputFile));
 		Boolean isSparkSessionManaged = isSparkSessionManaged(parser);
 		log.info("isSparkSessionManaged: {}", isSparkSessionManaged);

 		final String workingPath = parser.get("workingPath");
 		log.info("workingPath {}", workingPath);

+		final String format = parser.get("format");
+		log.info("format {}", format);
+
 		SparkConf sconf = new SparkConf();

 		final String delimiter = Optional
@ -64,16 +68,17 @@ public class ReadCOCI implements Serializable {
 					workingPath,
 					inputFile,
 					outputPath,
-					delimiter);
+					delimiter,
+					format);
 			});
 	}

 	private static void doRead(SparkSession spark, String workingPath, String[] inputFiles,
 		String outputPath,
-		String delimiter) throws IOException {
+		String delimiter, String format) {

 		for (String inputFile : inputFiles) {
-			String p_string = workingPath + "/" + inputFile + ".gz";
+			String pString = workingPath + "/" + inputFile + ".gz";

 			Dataset<Row> cociData = spark
 				.read()
@ -82,14 +87,20 @@ public class ReadCOCI implements Serializable {
 				.option("inferSchema", "true")
 				.option("header", "true")
 				.option("quotes", "\"")
-				.load(p_string)
+				.load(pString)
 				.repartition(100);

 			cociData.map((MapFunction<Row, COCI>) row -> {
 				COCI coci = new COCI();
+				if (format.equals("COCI")) {
+					coci.setCiting(row.getString(1));
+					coci.setCited(row.getString(2));
+				} else {
+					coci.setCiting(String.valueOf(row.getInt(1)));
+					coci.setCited(String.valueOf(row.getInt(2)));
+				}
 				coci.setOci(row.getString(0));
-				coci.setCiting(row.getString(1));
-				coci.setCited(row.getString(2));
+
 				return coci;
 			}, Encoders.bean(COCI.class))
 				.write()
--- a/dhp-workflows/dhp-aggregation/src/main/resources/eu/dnetlib/dhp/actionmanager/createunresolvedentities/oozie_app/workflow.xml
+++ b/dhp-workflows/dhp-aggregation/src/main/resources/eu/dnetlib/dhp/actionmanager/createunresolvedentities/oozie_app/workflow.xml
@ -5,11 +5,6 @@
            <name>fosPath</name>
            <description>the input path of the resources to be extended</description>
        </property>
-
-        <property>
-            <name>bipScorePath</name>
-            <description>the path where to find the bipFinder scores</description>
-        </property>
        <property>
            <name>outputPath</name>
            <description>the path where to store the actionset</description>
@ -77,35 +72,10 @@


    <fork name="prepareInfo">
-        <path start="prepareBip"/>
        <path start="getFOS"/>
        <path start="getSDG"/>
    </fork>

-    <action name="prepareBip">
-        <spark xmlns="uri:oozie:spark-action:0.2">
-            <master>yarn</master>
-            <mode>cluster</mode>
-            <name>Produces the unresolved from BIP! Finder</name>
-            <class>eu.dnetlib.dhp.actionmanager.createunresolvedentities.PrepareBipFinder</class>
-            <jar>dhp-aggregation-${projectVersion}.jar</jar>
-            <spark-opts>
-                --executor-memory=${sparkExecutorMemory}
-                --executor-cores=${sparkExecutorCores}
-                --driver-memory=${sparkDriverMemory}
-                --conf spark.extraListeners=${spark2ExtraListeners}
-                --conf spark.sql.queryExecutionListeners=${spark2SqlQueryExecutionListeners}
-                --conf spark.yarn.historyServer.address=${spark2YarnHistoryServerAddress}
-                --conf spark.eventLog.dir=${nameNode}${spark2EventLogDir}
-                --conf spark.sql.warehouse.dir=${sparkSqlWarehouseDir}
-            </spark-opts>
-            <arg>--sourcePath</arg><arg>${bipScorePath}</arg>
-            <arg>--outputPath</arg><arg>${workingDir}/prepared</arg>
-        </spark>
-        <ok to="join"/>
-        <error to="Kill"/>
-    </action>
-
    <action name="getFOS">
        <spark xmlns="uri:oozie:spark-action:0.2">
            <master>yarn</master>
@ -125,6 +95,7 @@
            </spark-opts>
            <arg>--sourcePath</arg><arg>${fosPath}</arg>
            <arg>--outputPath</arg><arg>${workingDir}/input/fos</arg>
+            <arg>--delimiter</arg><arg>${delimiter}</arg>
        </spark>
        <ok to="prepareFos"/>
        <error to="Kill"/>
@ -213,7 +184,7 @@
        <spark xmlns="uri:oozie:spark-action:0.2">
            <master>yarn</master>
            <mode>cluster</mode>
-            <name>Saves the result produced for bip and fos by grouping results with the same id</name>
+            <name>Save the unresolved entities grouping results with the same id</name>
            <class>eu.dnetlib.dhp.actionmanager.createunresolvedentities.SparkSaveUnresolved</class>
            <jar>dhp-aggregation-${projectVersion}.jar</jar>
            <spark-opts>
--- a/dhp-workflows/dhp-aggregation/src/main/resources/eu/dnetlib/dhp/actionmanager/opencitations/as_parameters.json
+++ b/dhp-workflows/dhp-aggregation/src/main/resources/eu/dnetlib/dhp/actionmanager/opencitations/as_parameters.json
@ -16,10 +16,11 @@
    "paramLongName": "isSparkSessionManaged",
    "paramDescription": "the hdfs name node",
    "paramRequired": false
-  },  {
-  "paramName": "sdr",
-  "paramLongName": "shouldDuplicateRels",
-  "paramDescription": "the hdfs name node",
-  "paramRequired": false
-}
+  },
+  {
+    "paramName": "sdr",
+    "paramLongName": "shouldDuplicateRels",
+    "paramDescription": "activates/deactivates the construction of bidirectional relations Cites/IsCitedBy",
+    "paramRequired": false
+  }
 ]
--- a/dhp-workflows/dhp-aggregation/src/main/resources/eu/dnetlib/dhp/actionmanager/opencitations/input_parameters.json
+++ b/dhp-workflows/dhp-aggregation/src/main/resources/eu/dnetlib/dhp/actionmanager/opencitations/input_parameters.json
@ -16,5 +16,11 @@
    "paramLongName": "hdfsNameNode",
    "paramDescription": "the hdfs name node",
    "paramRequired": true
+  },
+  {
+    "paramName": "p",
+    "paramLongName": "prefix",
+    "paramDescription": "COCI or POCI",
+    "paramRequired": true
  }
 ]
--- a/dhp-workflows/dhp-aggregation/src/main/resources/eu/dnetlib/dhp/actionmanager/opencitations/input_readcoci_parameters.json
+++ b/dhp-workflows/dhp-aggregation/src/main/resources/eu/dnetlib/dhp/actionmanager/opencitations/input_readcoci_parameters.json
@ -30,7 +30,12 @@
    "paramLongName": "inputFile",
    "paramDescription": "the hdfs name node",
    "paramRequired": true
-  }
+  }, {
+  "paramName": "f",
+  "paramLongName": "format",
+  "paramDescription": "the hdfs name node",
+  "paramRequired": true
+}
 ]


--- a/dhp-workflows/dhp-aggregation/src/main/resources/eu/dnetlib/dhp/actionmanager/opencitations/oozie_app/workflow.xml
+++ b/dhp-workflows/dhp-aggregation/src/main/resources/eu/dnetlib/dhp/actionmanager/opencitations/oozie_app/workflow.xml
@ -34,6 +34,7 @@
    <kill name="Kill">
        <message>Action failed, error message[${wf:errorMessage(wf:lastErrorNode())}]</message>
    </kill>
+
    <action name="download">
        <shell xmlns="uri:oozie:shell-action:0.2">
            <job-tracker>${jobTracker}</job-tracker>
@ -46,7 +47,7 @@
            </configuration>
            <exec>download.sh</exec>
            <argument>${filelist}</argument>
-            <argument>${workingPath}/Original</argument>
+            <argument>${workingPath}/${prefix}/Original</argument>
            <env-var>HADOOP_USER_NAME=${wf:user()}</env-var>
            <file>download.sh</file>
            <capture-output/>
@ -54,12 +55,14 @@
        <ok to="extract"/>
        <error to="Kill"/>
    </action>
+
    <action name="extract">
        <java>
            <main-class>eu.dnetlib.dhp.actionmanager.opencitations.GetOpenCitationsRefs</main-class>
            <arg>--hdfsNameNode</arg><arg>${nameNode}</arg>
            <arg>--inputFile</arg><arg>${inputFile}</arg>
-            <arg>--workingPath</arg><arg>${workingPath}</arg>
+            <arg>--workingPath</arg><arg>${workingPath}/${prefix}</arg>
+            <arg>--prefix</arg><arg>${prefix}</arg>
        </java>
        <ok to="read"/>
        <error to="Kill"/>
@ -82,10 +85,11 @@
                --conf spark.eventLog.dir=${nameNode}${spark2EventLogDir}
                --conf spark.sql.warehouse.dir=${sparkSqlWarehouseDir}
            </spark-opts>
-            <arg>--workingPath</arg><arg>${workingPath}/COCI</arg>
-            <arg>--outputPath</arg><arg>${workingPath}/COCI_JSON/</arg>
+            <arg>--workingPath</arg><arg>${workingPath}/${prefix}/${prefix}</arg>
+            <arg>--outputPath</arg><arg>${workingPath}/${prefix}/${prefix}_JSON/</arg>
            <arg>--delimiter</arg><arg>${delimiter}</arg>
            <arg>--inputFile</arg><arg>${inputFileCoci}</arg>
+            <arg>--format</arg><arg>${prefix}</arg>
        </spark>
        <ok to="create_actionset"/>
        <error to="Kill"/>
@ -108,7 +112,7 @@
                --conf spark.eventLog.dir=${nameNode}${spark2EventLogDir}
                --conf spark.sql.warehouse.dir=${sparkSqlWarehouseDir}
            </spark-opts>
-            <arg>--inputPath</arg><arg>${workingPath}/COCI_JSON</arg>
+            <arg>--inputPath</arg><arg>${workingPath}</arg>
            <arg>--outputPath</arg><arg>${outputPath}</arg>
        </spark>
        <ok to="End"/>
--- a/dhp-workflows/dhp-aggregation/src/main/resources/eu/dnetlib/dhp/datacite/hostedBy_map.json
+++ b/dhp-workflows/dhp-aggregation/src/main/resources/eu/dnetlib/dhp/datacite/hostedBy_map.json
@ -1,4 +1,9 @@
 {
+ "ETHZ.UNIGENF": {
+  "openaire_id": "opendoar____::1400",
+  "datacite_name": "Uni Genf",
+  "official_name": "Archive ouverte UNIGE"
+ },
 "GESIS.RKI": {
  "openaire_id": "re3data_____::r3d100010436",
  "datacite_name": "Forschungsdatenzentrum  am Robert Koch Institut",
--- a/dhp-workflows/dhp-aggregation/src/test/java/eu/dnetlib/dhp/actionmanager/createunresolvedentities/GetFosTest.java
+++ b/dhp-workflows/dhp-aggregation/src/test/java/eu/dnetlib/dhp/actionmanager/createunresolvedentities/GetFosTest.java
@ -13,10 +13,7 @@ import org.apache.spark.SparkConf;
 import org.apache.spark.api.java.JavaRDD;
 import org.apache.spark.api.java.JavaSparkContext;
 import org.apache.spark.sql.SparkSession;
-import org.junit.jupiter.api.AfterAll;
-import org.junit.jupiter.api.Assertions;
-import org.junit.jupiter.api.BeforeAll;
-import org.junit.jupiter.api.Test;
+import org.junit.jupiter.api.*;
 import org.slf4j.Logger;
 import org.slf4j.LoggerFactory;

@ -68,6 +65,7 @@ public class GetFosTest {
 	}

 	@Test
+	@Disabled
 	void test3() throws Exception {
 		final String sourcePath = getClass()
 			.getResource("/eu/dnetlib/dhp/actionmanager/createunresolvedentities/fos/fos_sbs.tsv")
@ -96,4 +94,37 @@ public class GetFosTest {
 		tmp.foreach(t -> Assertions.assertTrue(t.getLevel3() != null));

 	}
+
+	@Test
+	void test4() throws Exception {
+		final String sourcePath = getClass()
+			.getResource("/eu/dnetlib/dhp/actionmanager/createunresolvedentities/fos/fos_sbs2.csv")
+			.getPath();
+
+		final String outputPath = workingDir.toString() + "/fos.json";
+		GetFOSSparkJob
+			.main(
+				new String[] {
+					"--isSparkSessionManaged", Boolean.FALSE.toString(),
+					"--sourcePath", sourcePath,
+					"--delimiter", ",",
+					"-outputPath", outputPath
+
+				});
+
+		final JavaSparkContext sc = JavaSparkContext.fromSparkContext(spark.sparkContext());
+
+		JavaRDD<FOSDataModel> tmp = sc
+			.textFile(outputPath)
+			.map(item -> OBJECT_MAPPER.readValue(item, FOSDataModel.class));
+
+		tmp.foreach(t -> Assertions.assertTrue(t.getDoi() != null));
+		tmp.foreach(t -> Assertions.assertTrue(t.getLevel1() != null));
+		tmp.foreach(t -> Assertions.assertTrue(t.getLevel2() != null));
+		tmp.foreach(t -> Assertions.assertTrue(t.getLevel3() != null));
+		tmp.foreach(t -> Assertions.assertTrue(t.getLevel4() != null));
+		tmp.foreach(t -> Assertions.assertTrue(t.getScoreL3() != null));
+		tmp.foreach(t -> Assertions.assertTrue(t.getScoreL4() != null));
+
+	}
 }
--- a/dhp-workflows/dhp-aggregation/src/test/java/eu/dnetlib/dhp/actionmanager/createunresolvedentities/PrepareTest.java
+++ b/dhp-workflows/dhp-aggregation/src/test/java/eu/dnetlib/dhp/actionmanager/createunresolvedentities/PrepareTest.java
@ -67,92 +67,6 @@ public class PrepareTest {
 		spark.stop();
 	}

-	@Test
-	void bipPrepareTest() throws Exception {
-		final String sourcePath = getClass()
-			.getResource("/eu/dnetlib/dhp/actionmanager/createunresolvedentities/bip/bip.json")
-			.getPath();
-
-		PrepareBipFinder
-			.main(
-				new String[] {
-					"--isSparkSessionManaged", Boolean.FALSE.toString(),
-					"--sourcePath", sourcePath,
-					"--outputPath", workingDir.toString() + "/work"
-
-				});
-
-		final JavaSparkContext sc = JavaSparkContext.fromSparkContext(spark.sparkContext());
-
-		JavaRDD<Result> tmp = sc
-			.textFile(workingDir.toString() + "/work/bip")
-			.map(item -> OBJECT_MAPPER.readValue(item, Result.class));
-
-		Assertions.assertEquals(86, tmp.count());
-
-		String doi1 = "unresolved::10.0000/096020199389707::doi";
-
-		Assertions.assertEquals(1, tmp.filter(r -> r.getId().equals(doi1)).count());
-		Assertions.assertEquals(1, tmp.filter(r -> r.getId().equals(doi1)).collect().get(0).getInstance().size());
-		Assertions
-			.assertEquals(
-				3, tmp.filter(r -> r.getId().equals(doi1)).collect().get(0).getInstance().get(0).getMeasures().size());
-		Assertions
-			.assertEquals(
-				"6.34596412687e-09", tmp
-					.filter(r -> r.getId().equals(doi1))
-					.collect()
-					.get(0)
-					.getInstance()
-					.get(0)
-					.getMeasures()
-					.stream()
-					.filter(sl -> sl.getId().equals("influence"))
-					.collect(Collectors.toList())
-					.get(0)
-					.getUnit()
-					.get(0)
-					.getValue());
-		Assertions
-			.assertEquals(
-				"0.641151896994", tmp
-					.filter(r -> r.getId().equals(doi1))
-					.collect()
-					.get(0)
-					.getInstance()
-					.get(0)
-					.getMeasures()
-					.stream()
-					.filter(sl -> sl.getId().equals("popularity_alt"))
-					.collect(Collectors.toList())
-					.get(0)
-					.getUnit()
-					.get(0)
-					.getValue());
-		Assertions
-			.assertEquals(
-				"2.33375102921e-09", tmp
-					.filter(r -> r.getId().equals(doi1))
-					.collect()
-					.get(0)
-					.getInstance()
-					.get(0)
-					.getMeasures()
-					.stream()
-					.filter(sl -> sl.getId().equals("popularity"))
-					.collect(Collectors.toList())
-					.get(0)
-					.getUnit()
-					.get(0)
-					.getValue());
-
-		final String doi2 = "unresolved::10.3390/s18072310::doi";
-
-		Assertions.assertEquals(1, tmp.filter(r -> r.getId().equals(doi2)).count());
-		Assertions.assertEquals(1, tmp.filter(r -> r.getId().equals(doi2)).collect().get(0).getInstance().size());
-
-	}
-
 	@Test
 	void fosPrepareTest() throws Exception {
 		final String sourcePath = getClass()
@ -222,6 +136,76 @@ public class PrepareTest {

 	}

+	@Test
+	void fosPrepareTest2() throws Exception {
+		final String sourcePath = getClass()
+			.getResource("/eu/dnetlib/dhp/actionmanager/createunresolvedentities/fos/fos_sbs_2.json")
+			.getPath();
+
+		PrepareFOSSparkJob
+			.main(
+				new String[] {
+					"--isSparkSessionManaged", Boolean.FALSE.toString(),
+					"--sourcePath", sourcePath,
+
+					"-outputPath", workingDir.toString() + "/work"
+
+				});
+
+		final JavaSparkContext sc = JavaSparkContext.fromSparkContext(spark.sparkContext());
+
+		JavaRDD<Result> tmp = sc
+			.textFile(workingDir.toString() + "/work/fos")
+			.map(item -> OBJECT_MAPPER.readValue(item, Result.class));
+
+		String doi1 = "unresolved::10.1016/j.revmed.2006.07.012::doi";
+
+		assertEquals(13, tmp.count());
+		assertEquals(1, tmp.filter(row -> row.getId().equals(doi1)).count());
+
+		Result result = tmp
+			.filter(r -> r.getId().equals(doi1))
+			.first();
+
+		result.getSubject().forEach(s -> System.out.println(s.getValue() + " trust = " + s.getDataInfo().getTrust()));
+		Assertions.assertEquals(6, result.getSubject().size());
+
+		assertTrue(
+			result
+				.getSubject()
+				.stream()
+				.anyMatch(
+					s -> s.getValue().contains("03 medical and health sciences")
+						&& s.getDataInfo().getTrust().equals("")));
+
+		assertTrue(
+			result
+				.getSubject()
+				.stream()
+				.anyMatch(
+					s -> s.getValue().contains("0302 clinical medicine") && s.getDataInfo().getTrust().equals("")));
+
+		assertTrue(
+			result
+				.getSubject()
+				.stream()
+				.anyMatch(
+					s -> s
+						.getValue()
+						.contains("030204 cardiovascular system & hematology")
+						&& s.getDataInfo().getTrust().equals("0.5101401805877686")));
+		assertTrue(
+			result
+				.getSubject()
+				.stream()
+				.anyMatch(
+					s -> s
+						.getValue()
+						.contains("03020409 Hematology/Coagulopathies")
+						&& s.getDataInfo().getTrust().equals("0.0546871414174914")));
+
+	}
+
 	@Test
 	void sdgPrepareTest() throws Exception {
 		final String sourcePath = getClass()
@ -268,57 +252,4 @@ public class PrepareTest {

 	}

-//	@Test
-//	void test3() throws Exception {
-//		final String sourcePath = "/Users/miriam.baglioni/Downloads/doi_fos_results_20_12_2021.csv.gz";
-//
-//		final String outputPath = workingDir.toString() + "/fos.json";
-//		GetFOSSparkJob
-//			.main(
-//				new String[] {
-//					"--isSparkSessionManaged", Boolean.FALSE.toString(),
-//					"--sourcePath", sourcePath,
-//
-//					"-outputPath", outputPath
-//
-//				});
-//
-//		final JavaSparkContext sc = JavaSparkContext.fromSparkContext(spark.sparkContext());
-//
-//		JavaRDD<FOSDataModel> tmp = sc
-//			.textFile(outputPath)
-//			.map(item -> OBJECT_MAPPER.readValue(item, FOSDataModel.class));
-//
-//		tmp.foreach(t -> Assertions.assertTrue(t.getDoi() != null));
-//		tmp.foreach(t -> Assertions.assertTrue(t.getLevel1() != null));
-//		tmp.foreach(t -> Assertions.assertTrue(t.getLevel2() != null));
-//		tmp.foreach(t -> Assertions.assertTrue(t.getLevel3() != null));
-//
-//	}
-//
-//	@Test
-//	void test4() throws Exception {
-//		final String sourcePath = "/Users/miriam.baglioni/Downloads/doi_sdg_results_20_12_21.csv.gz";
-//
-//		final String outputPath = workingDir.toString() + "/sdg.json";
-//		GetSDGSparkJob
-//			.main(
-//				new String[] {
-//					"--isSparkSessionManaged", Boolean.FALSE.toString(),
-//					"--sourcePath", sourcePath,
-//
-//					"-outputPath", outputPath
-//
-//				});
-//
-//		final JavaSparkContext sc = JavaSparkContext.fromSparkContext(spark.sparkContext());
-//
-//		JavaRDD<SDGDataModel> tmp = sc
-//			.textFile(outputPath)
-//			.map(item -> OBJECT_MAPPER.readValue(item, SDGDataModel.class));
-//
-//		tmp.foreach(t -> Assertions.assertTrue(t.getDoi() != null));
-//		tmp.foreach(t -> Assertions.assertTrue(t.getSbj() != null));
-//
-//	}
 }
--- a/dhp-workflows/dhp-aggregation/src/test/java/eu/dnetlib/dhp/actionmanager/createunresolvedentities/ProduceTest.java
+++ b/dhp-workflows/dhp-aggregation/src/test/java/eu/dnetlib/dhp/actionmanager/createunresolvedentities/ProduceTest.java
@ -340,18 +340,7 @@ public class ProduceTest {
 	}

 	private JavaRDD<Result> getResultJavaRDD() throws Exception {
-		final String bipPath = getClass()
-			.getResource("/eu/dnetlib/dhp/actionmanager/createunresolvedentities/bip/bip.json")
-			.getPath();

-		PrepareBipFinder
-			.main(
-				new String[] {
-					"--isSparkSessionManaged", Boolean.FALSE.toString(),
-					"--sourcePath", bipPath,
-					"--outputPath", workingDir.toString() + "/work"
-
-				});
 		final String fosPath = getClass()
 			.getResource("/eu/dnetlib/dhp/actionmanager/createunresolvedentities/fos/fos.json")
 			.getPath();
@ -379,6 +368,40 @@ public class ProduceTest {
 			.map(item -> OBJECT_MAPPER.readValue(item, Result.class));
 	}

+	@Test
+	public JavaRDD<Result> getResultFosJavaRDD() throws Exception {
+
+		final String fosPath = getClass()
+			.getResource("/eu/dnetlib/dhp/actionmanager/createunresolvedentities/fos/fos_sbs_2.json")
+			.getPath();
+
+		PrepareFOSSparkJob
+			.main(
+				new String[] {
+					"--isSparkSessionManaged", Boolean.FALSE.toString(),
+					"--sourcePath", fosPath,
+					"-outputPath", workingDir.toString() + "/work"
+				});
+
+		SparkSaveUnresolved.main(new String[] {
+			"--isSparkSessionManaged", Boolean.FALSE.toString(),
+			"--sourcePath", workingDir.toString() + "/work",
+
+			"-outputPath", workingDir.toString() + "/unresolved"
+
+		});
+
+		final JavaSparkContext sc = JavaSparkContext.fromSparkContext(spark.sparkContext());
+
+		JavaRDD<Result> tmp = sc
+			.textFile(workingDir.toString() + "/unresolved")
+			.map(item -> OBJECT_MAPPER.readValue(item, Result.class));
+		tmp.foreach(r -> System.out.println(new ObjectMapper().writeValueAsString(r)));
+
+		return tmp;
+
+	}
+
 	@Test
 	void prepareTest5Subjects() throws Exception {
 		final String doi = "unresolved::10.1063/5.0032658::doi";
@ -415,18 +438,7 @@ public class ProduceTest {
 	}

 	private JavaRDD<Result> getResultJavaRDDPlusSDG() throws Exception {
-		final String bipPath = getClass()
-			.getResource("/eu/dnetlib/dhp/actionmanager/createunresolvedentities/bip/bip.json")
-			.getPath();

-		PrepareBipFinder
-			.main(
-				new String[] {
-					"--isSparkSessionManaged", Boolean.FALSE.toString(),
-					"--sourcePath", bipPath,
-					"--outputPath", workingDir.toString() + "/work"
-
-				});
 		final String fosPath = getClass()
 			.getResource("/eu/dnetlib/dhp/actionmanager/createunresolvedentities/fos/fos.json")
 			.getPath();
@ -483,14 +495,6 @@ public class ProduceTest {
 					.filter(row -> row.getSubject() != null)
 					.count());

-		Assertions
-			.assertEquals(
-				85,
-				tmp
-					.filter(row -> !row.getId().equals(doi))
-					.filter(r -> r.getInstance() != null && r.getInstance().size() > 0)
-					.count());
-
 	}

 	@Test
--- a/dhp-workflows/dhp-aggregation/src/test/resources/eu/dnetlib/dhp/actionmanager/createunresolvedentities/fos/fos_sbs2.csv
+++ b/dhp-workflows/dhp-aggregation/src/test/resources/eu/dnetlib/dhp/actionmanager/createunresolvedentities/fos/fos_sbs2.csv
@ -0,0 +1,26 @@
+DOI,OAID,level1,level2,level3,level4,score_for_L3,score_for_L4
+10.1016/j.anucene.2006.02.004,doi_________::00059d9963edf633bec756fb21b5bd72,02 engineering and technology,"0202 electrical engineering, electronic engineering, information engineering",020209 energy,02020908 Climate change policy/Ethanol fuel,0.5,0.5
+10.1016/j.anucene.2006.02.004,doi_________::00059d9963edf633bec756fb21b5bd72,02 engineering and technology,0211 other engineering and technologies,021108 energy,02110808 Climate change policy/Ethanol fuel,0.5,0.5
+10.1016/j.revmed.2006.07.010,doi_________::0026476c1651a92c933d752ff12496c7,03 medical and health sciences,0302 clinical medicine,030220 oncology & carcinogenesis,N/A,0.5036656856536865,0.0
+10.1016/j.revmed.2006.07.010,doi_________::0026476c1651a92c933d752ff12496c7,03 medical and health sciences,0302 clinical medicine,030212 general & internal medicine,N/A,0.4963343143463135,0.0
+10.20965/jrm.2006.p0312,doi_________::0028336a2f3826cc83c47dbefac71543,02 engineering and technology,0209 industrial biotechnology,020901 industrial engineering & automation,02090104 Robotics/Robots,0.6111094951629639,0.5053805979936855
+10.20965/jrm.2006.p0312,doi_________::0028336a2f3826cc83c47dbefac71543,01 natural sciences,0104 chemical sciences,010401 analytical chemistry,N/A,0.3888905048370361,0.0
+10.1111/j.1747-7379.2006.040_1.x,doi_________::002c7077e7c114a8304eb90f59e45fa4,05 social sciences,0506 political science,050602 political science & public administration,05060202 Ethnic groups/Ethnicity,0.6159052848815918,0.7369035568037298
+10.1111/j.1747-7379.2006.040_1.x,doi_________::002c7077e7c114a8304eb90f59e45fa4,05 social sciences,0502 economics and business,050207 economics,N/A,0.3840946555137634,0.0
+10.1007/s10512-006-0049-9,doi_________::003f29f9254819cf4c78558b1bc25f10,02 engineering and technology,"0202 electrical engineering, electronic engineering, information engineering",020209 energy,02020908 Climate change policy/Ethanol fuel,0.5,0.5
+10.1007/s10512-006-0049-9,doi_________::003f29f9254819cf4c78558b1bc25f10,02 engineering and technology,0211 other engineering and technologies,021108 energy,02110808 Climate change policy/Ethanol fuel,0.5,0.5
+10.1111/j.1365-2621.2005.01045.x,doi_________::00419355b4c3e0646bd0e1b301164c8e,04 agricultural and veterinary sciences,0404 agricultural biotechnology,040401 food science,04040102 Food science/Food industry,0.5,0.5
+10.1111/j.1365-2621.2005.01045.x,doi_________::00419355b4c3e0646bd0e1b301164c8e,04 agricultural and veterinary sciences,0405 other agricultural sciences,040502 food science,04050202 Food science/Food industry,0.5,0.5
+10.1002/chin.200617262,doi_________::004c8cef80668904961b9e62841793c8,01 natural sciences,0104 chemical sciences,010405 organic chemistry,01040508 Functional groups/Ethers,0.5566747188568115,0.5582916736602783
+10.1002/chin.200617262,doi_________::004c8cef80668904961b9e62841793c8,01 natural sciences,0104 chemical sciences,010402 general chemistry,01040207 Chemical synthesis/Total synthesis,0.4433253407478332,0.4417082965373993
+10.1016/j.revmed.2006.07.012,doi_________::005b1d0fb650b680abaf6cfe26a21604,03 medical and health sciences,0302 clinical medicine,030204 cardiovascular system & hematology,03020409 Hematology/Coagulopathies,0.5101401805877686,0.0546871414174914
+10.1016/j.revmed.2006.07.012,doi_________::005b1d0fb650b680abaf6cfe26a21604,03 medical and health sciences,0301 basic medicine,030105 genetics & heredity,N/A,0.4898599088191986,0.0
+10.4109/jslab.17.132,doi_________::00889baa06de363e37930daaf8e800c0,03 medical and health sciences,0301 basic medicine,030104 developmental biology,N/A,0.5,0.0
+10.4109/jslab.17.132,doi_________::00889baa06de363e37930daaf8e800c0,03 medical and health sciences,0303 health sciences,030304 developmental biology,N/A,0.5,0.0
+10.1108/00251740610715687,doi_________::0092cb1b1920d556719385a26363ecaa,05 social sciences,0502 economics and business,050203 business & management,05020311 International business/International trade,0.605047881603241,0.2156608108845153
+10.1108/00251740610715687,doi_________::0092cb1b1920d556719385a26363ecaa,05 social sciences,0502 economics and business,050211 marketing,N/A,0.394952118396759,0.0
+10.1080/03067310500248098,doi_________::00a76678d230e3f20b6356804448028f,04 agricultural and veterinary sciences,0404 agricultural biotechnology,040401 food science,04040102 Food science/Food industry,0.5,0.5
+10.1080/03067310500248098,doi_________::00a76678d230e3f20b6356804448028f,04 agricultural and veterinary sciences,0405 other agricultural sciences,040502 food science,04050202 Food science/Food industry,0.5,0.5
+10.3152/147154306781778533,doi_________::00acc520f3939e5a6675343881fed4f2,05 social sciences,0502 economics and business,050203 business & management,05020307 Innovation/Product management,0.5293408632278442,0.5326762795448303
+10.3152/147154306781778533,doi_________::00acc520f3939e5a6675343881fed4f2,05 social sciences,0509 other social sciences,050905 science studies,05090502 Social philosophy/Capitalism,0.4706590473651886,0.4673237204551697
+10.1785/0120050806,doi_________::00d5831d329e7ae4523d78bfc3042e98,02 engineering and technology,0211 other engineering and technologies,021101 geological & geomatics engineering,02110103 Concrete/Building materials,0.5343400835990906,0.3285667930180677
--- a/dhp-workflows/dhp-aggregation/src/test/resources/eu/dnetlib/dhp/actionmanager/createunresolvedentities/fos/fos_sbs_2.json
+++ b/dhp-workflows/dhp-aggregation/src/test/resources/eu/dnetlib/dhp/actionmanager/createunresolvedentities/fos/fos_sbs_2.json
@ -0,0 +1,25 @@
+{"doi":"10.1016/j.anucene.2006.02.004","level1":"02 engineering and technology","level2":"0202 electrical engineering, electronic engineering, information engineering","level3":"020209 energy","level4":"02020908 Climate change policy/Ethanol fuel","scoreL3":"0.5","scoreL4":"0.5"}
+{"doi":"10.1016/j.anucene.2006.02.004","level1":"02 engineering and technology","level2":"0211 other engineering and technologies","level3":"021108 energy","level4":"02110808 Climate change policy/Ethanol fuel","scoreL3":"0.5","scoreL4":"0.5"}
+{"doi":"10.1016/j.revmed.2006.07.010","level1":"03 medical and health sciences","level2":"0302 clinical medicine","level3":"030220 oncology & carcinogenesis","level4":"N/A","scoreL3":"0.5036656856536865","scoreL4":"0.0"}
+{"doi":"10.1016/j.revmed.2006.07.010","level1":"03 medical and health sciences","level2":"0302 clinical medicine","level3":"030212 general & internal medicine","level4":"N/A","scoreL3":"0.4963343143463135","scoreL4":"0.0"}
+{"doi":"10.20965/jrm.2006.p0312","level1":"02 engineering and technology","level2":"0209 industrial biotechnology","level3":"020901 industrial engineering & automation","level4":"02090104 Robotics/Robots","scoreL3":"0.6111094951629639","scoreL4":"0.5053805979936855"}
+{"doi":"10.20965/jrm.2006.p0312","level1":"01 natural sciences","level2":"0104 chemical sciences","level3":"010401 analytical chemistry","level4":"N/A","scoreL3":"0.3888905048370361","scoreL4":"0.0"}
+{"doi":"10.1111/j.1747-7379.2006.040_1.x","level1":"05 social sciences","level2":"0506 political science","level3":"050602 political science & public administration","level4":"05060202 Ethnic groups/Ethnicity","scoreL3":"0.6159052848815918","scoreL4":"0.7369035568037298"}
+{"doi":"10.1111/j.1747-7379.2006.040_1.x","level1":"05 social sciences","level2":"0502 economics and business","level3":"050207 economics","level4":"N/A","scoreL3":"0.3840946555137634","scoreL4":"0.0"}
+{"doi":"10.1007/s10512-006-0049-9","level1":"02 engineering and technology","level2":"0202 electrical engineering, electronic engineering, information engineering","level3":"020209 energy","level4":"02020908 Climate change policy/Ethanol fuel","scoreL3":"0.5","scoreL4":"0.5"}
+{"doi":"10.1007/s10512-006-0049-9","level1":"02 engineering and technology","level2":"0211 other engineering and technologies","level3":"021108 energy","level4":"02110808 Climate change policy/Ethanol fuel","scoreL3":"0.5","scoreL4":"0.5"}
+{"doi":"10.1111/j.1365-2621.2005.01045.x","level1":"04 agricultural and veterinary sciences","level2":"0404 agricultural biotechnology","level3":"040401 food science","level4":"04040102 Food science/Food industry","scoreL3":"0.5","scoreL4":"0.5"}
+{"doi":"10.1111/j.1365-2621.2005.01045.x","level1":"04 agricultural and veterinary sciences","level2":"0405 other agricultural sciences","level3":"040502 food science","level4":"04050202 Food science/Food industry","scoreL3":"0.5","scoreL4":"0.5"}
+{"doi":"10.1002/chin.200617262","level1":"01 natural sciences","level2":"0104 chemical sciences","level3":"010405 organic chemistry","level4":"01040508 Functional groups/Ethers","scoreL3":"0.5566747188568115","scoreL4":"0.5582916736602783"}
+{"doi":"10.1002/chin.200617262","level1":"01 natural sciences","level2":"0104 chemical sciences","level3":"010402 general chemistry","level4":"01040207 Chemical synthesis/Total synthesis","scoreL3":"0.4433253407478332","scoreL4":"0.4417082965373993"}
+{"doi":"10.1016/j.revmed.2006.07.012","level1":"03 medical and health sciences","level2":"0302 clinical medicine","level3":"030204 cardiovascular system & hematology","level4":"03020409 Hematology/Coagulopathies","scoreL3":"0.5101401805877686","scoreL4":"0.0546871414174914"}
+{"doi":"10.1016/j.revmed.2006.07.012","level1":"03 medical and health sciences","level2":"0301 basic medicine","level3":"030105 genetics & heredity","level4":"N/A","scoreL3":"0.4898599088191986","scoreL4":"0.0"}
+{"doi":"10.4109/jslab.17.132","level1":"03 medical and health sciences","level2":"0301 basic medicine","level3":"030104 developmental biology","level4":"N/A","scoreL3":"0.5","scoreL4":"0.0"}
+{"doi":"10.4109/jslab.17.132","level1":"03 medical and health sciences","level2":"0303 health sciences","level3":"030304 developmental biology","level4":"N/A","scoreL3":"0.5","scoreL4":"0.0"}
+{"doi":"10.1108/00251740610715687","level1":"05 social sciences","level2":"0502 economics and business","level3":"050203 business & management","level4":"05020311 International business/International trade","scoreL3":"0.605047881603241","scoreL4":"0.2156608108845153"}
+{"doi":"10.1108/00251740610715687","level1":"05 social sciences","level2":"0502 economics and business","level3":"050211 marketing","level4":"N/A","scoreL3":"0.394952118396759","scoreL4":"0.0"}
+{"doi":"10.1080/03067310500248098","level1":"04 agricultural and veterinary sciences","level2":"0404 agricultural biotechnology","level3":"040401 food science","level4":"04040102 Food science/Food industry","scoreL3":"0.5","scoreL4":"0.5"}
+{"doi":"10.1080/03067310500248098","level1":"04 agricultural and veterinary sciences","level2":"0405 other agricultural sciences","level3":"040502 food science","level4":"04050202 Food science/Food industry","scoreL3":"0.5","scoreL4":"0.5"}
+{"doi":"10.3152/147154306781778533","level1":"05 social sciences","level2":"0502 economics and business","level3":"050203 business & management","level4":"05020307 Innovation/Product management","scoreL3":"0.5293408632278442","scoreL4":"0.5326762795448303"}
+{"doi":"10.3152/147154306781778533","level1":"05 social sciences","level2":"0509 other social sciences","level3":"050905 science studies","level4":"05090502 Social philosophy/Capitalism","scoreL3":"0.4706590473651886","scoreL4":"0.4673237204551697"}
+{"doi":"10.1785/0120050806","level1":"02 engineering and technology","level2":"0211 other engineering and technologies","level3":"021101 geological & geomatics engineering","level4":"02110103 Concrete/Building materials","scoreL3":"0.5343400835990906","scoreL4":"0.3285667930180677"}
--- a/dhp-workflows/dhp-broker-events/src/main/java/eu/dnetlib/dhp/broker/oa/util/TrustUtils.java
+++ b/dhp-workflows/dhp-broker-events/src/main/java/eu/dnetlib/dhp/broker/oa/util/TrustUtils.java
@ -2,7 +2,9 @@
 package eu.dnetlib.dhp.broker.oa.util;

 import java.io.IOException;
+import java.nio.charset.StandardCharsets;

+import org.apache.commons.io.IOUtils;
 import org.apache.spark.sql.Row;
 import org.slf4j.Logger;
 import org.slf4j.LoggerFactory;
@ -27,10 +29,14 @@ public class TrustUtils {
 	static {
 		mapper = new ObjectMapper();
 		try {
-			dedupConfig = mapper
-				.readValue(
-					DedupConfig.class.getResourceAsStream("/eu/dnetlib/dhp/broker/oa/dedupConfig/dedupConfig.json"),
-					DedupConfig.class);
+			dedupConfig = DedupConfig
+				.load(
+					IOUtils
+						.toString(
+							DedupConfig.class
+								.getResourceAsStream("/eu/dnetlib/dhp/broker/oa/dedupConfig/dedupConfig.json"),
+							StandardCharsets.UTF_8));
+
 			deduper = new SparkDeduper(dedupConfig);
 		} catch (final IOException e) {
 			log.error("Error loading dedupConfig, e");
@ -57,7 +63,7 @@ public class TrustUtils {
 			return TrustUtils.rescale(score, threshold);
 		} catch (final Exception e) {
 			log.error("Error computing score between results", e);
-			return BrokerConstants.MIN_TRUST;
+			throw new RuntimeException(e);
 		}
 	}

--- a/dhp-workflows/dhp-dedup-openaire/src/main/java/eu/dnetlib/dhp/oa/dedup/RelationAggregator.java
+++ b/dhp-workflows/dhp-dedup-openaire/src/main/java/eu/dnetlib/dhp/oa/dedup/RelationAggregator.java
@ -1,57 +0,0 @@
-
-package eu.dnetlib.dhp.oa.dedup;
-
-import java.util.Objects;
-
-import org.apache.spark.sql.Encoder;
-import org.apache.spark.sql.Encoders;
-import org.apache.spark.sql.expressions.Aggregator;
-
-import eu.dnetlib.dhp.schema.oaf.Relation;
-
-public class RelationAggregator extends Aggregator<Relation, Relation, Relation> {
-
-	private static final Relation ZERO = new Relation();
-
-	@Override
-	public Relation zero() {
-		return ZERO;
-	}
-
-	@Override
-	public Relation reduce(Relation b, Relation a) {
-		return mergeRel(b, a);
-	}
-
-	@Override
-	public Relation merge(Relation b, Relation a) {
-		return mergeRel(b, a);
-	}
-
-	@Override
-	public Relation finish(Relation r) {
-		return r;
-	}
-
-	private Relation mergeRel(Relation b, Relation a) {
-		if (Objects.equals(b, ZERO)) {
-			return a;
-		}
-		if (Objects.equals(a, ZERO)) {
-			return b;
-		}
-
-		b.mergeFrom(a);
-		return b;
-	}
-
-	@Override
-	public Encoder<Relation> bufferEncoder() {
-		return Encoders.kryo(Relation.class);
-	}
-
-	@Override
-	public Encoder<Relation> outputEncoder() {
-		return Encoders.kryo(Relation.class);
-	}
-}
--- a/dhp-workflows/dhp-dedup-openaire/src/main/java/eu/dnetlib/dhp/oa/dedup/SparkCleanRelation.scala
+++ b/dhp-workflows/dhp-dedup-openaire/src/main/java/eu/dnetlib/dhp/oa/dedup/SparkCleanRelation.scala
@ -1,78 +0,0 @@
-package eu.dnetlib.dhp.oa.dedup
-
-import eu.dnetlib.dhp.application.ArgumentApplicationParser
-import eu.dnetlib.dhp.common.HdfsSupport
-import eu.dnetlib.dhp.schema.oaf.Relation
-import eu.dnetlib.dhp.utils.ISLookupClientFactory
-import eu.dnetlib.enabling.is.lookup.rmi.ISLookUpService
-import org.apache.commons.io.IOUtils
-import org.apache.spark.SparkConf
-import org.apache.spark.sql._
-import org.apache.spark.sql.functions.col
-import org.apache.spark.sql.types.{DataTypes, StructField, StructType}
-import org.slf4j.LoggerFactory
-
-object SparkCleanRelation {
-  private val log = LoggerFactory.getLogger(classOf[SparkCleanRelation])
-
-  @throws[Exception]
-  def main(args: Array[String]): Unit = {
-    val parser = new ArgumentApplicationParser(
-      IOUtils.toString(
-        classOf[SparkCleanRelation].getResourceAsStream("/eu/dnetlib/dhp/oa/dedup/cleanRelation_parameters.json")
-      )
-    )
-    parser.parseArgument(args)
-    val conf = new SparkConf
-
-    new SparkCleanRelation(parser, AbstractSparkAction.getSparkSession(conf))
-      .run(ISLookupClientFactory.getLookUpService(parser.get("isLookUpUrl")))
-  }
-}
-
-class SparkCleanRelation(parser: ArgumentApplicationParser, spark: SparkSession)
-    extends AbstractSparkAction(parser, spark) {
-  override def run(isLookUpService: ISLookUpService): Unit = {
-    val graphBasePath = parser.get("graphBasePath")
-    val inputPath = parser.get("inputPath")
-    val outputPath = parser.get("outputPath")
-
-    SparkCleanRelation.log.info("graphBasePath: '{}'", graphBasePath)
-    SparkCleanRelation.log.info("inputPath: '{}'", inputPath)
-    SparkCleanRelation.log.info("outputPath: '{}'", outputPath)
-
-    AbstractSparkAction.removeOutputDir(spark, outputPath)
-
-    val entities =
-      Seq("datasource", "project", "organization", "publication", "dataset", "software", "otherresearchproduct")
-
-    val idsSchema = StructType.fromDDL("`id` STRING, `dataInfo` STRUCT<`deletedbyinference`:BOOLEAN,`invisible`:BOOLEAN>")
-
-    val emptyIds = spark.createDataFrame(spark.sparkContext.emptyRDD[Row].setName("empty"),
-      idsSchema)
-
-    val ids = entities
-      .foldLeft(emptyIds)((ds, entity) => {
-        val entityPath = graphBasePath + '/' + entity
-        if (HdfsSupport.exists(entityPath, spark.sparkContext.hadoopConfiguration)) {
-          ds.union(spark.read.schema(idsSchema).json(entityPath))
-        } else {
-          ds
-        }
-      })
-      .filter("dataInfo.deletedbyinference != true AND dataInfo.invisible != true")
-      .select("id")
-      .distinct()
-
-    val relations = spark.read.schema(Encoders.bean(classOf[Relation]).schema).json(inputPath)
-      .filter("dataInfo.deletedbyinference != true AND dataInfo.invisible != true")
-
-    AbstractSparkAction.save(
-      relations
-        .join(ids, col("source") === ids("id"), "leftsemi")
-        .join(ids, col("target") === ids("id"), "leftsemi"),
-      outputPath,
-      SaveMode.Overwrite
-    )
-  }
-}
--- a/dhp-workflows/dhp-dedup-openaire/src/main/java/eu/dnetlib/dhp/oa/dedup/SparkCopyOpenorgsMergeRels.java
+++ b/dhp-workflows/dhp-dedup-openaire/src/main/java/eu/dnetlib/dhp/oa/dedup/SparkCopyOpenorgsMergeRels.java
@ -7,6 +7,7 @@ import java.util.Optional;
 import org.apache.commons.io.IOUtils;
 import org.apache.spark.SparkConf;
 import org.apache.spark.api.java.JavaRDD;
+import org.apache.spark.sql.Dataset;
 import org.apache.spark.sql.Encoders;
 import org.apache.spark.sql.SaveMode;
 import org.apache.spark.sql.SparkSession;
@ -77,13 +78,12 @@ public class SparkCopyOpenorgsMergeRels extends AbstractSparkAction {

 		log.info("Number of Openorgs Merge Relations collected: {}", mergeRelsRDD.count());

-		spark
+		final Dataset<Relation> relations = spark
 			.createDataset(
 				mergeRelsRDD.rdd(),
-				Encoders.bean(Relation.class))
-			.write()
-			.mode(SaveMode.Append)
-			.parquet(outputPath);
+				Encoders.bean(Relation.class));
+
+		saveParquet(relations, outputPath, SaveMode.Append);
 	}

 	private boolean isMergeRel(Relation rel) {
--- a/dhp-workflows/dhp-dedup-openaire/src/main/java/eu/dnetlib/dhp/oa/dedup/SparkCopyRelationsNoOpenorgs.java
+++ b/dhp-workflows/dhp-dedup-openaire/src/main/java/eu/dnetlib/dhp/oa/dedup/SparkCopyRelationsNoOpenorgs.java
@ -67,12 +67,7 @@ public class SparkCopyRelationsNoOpenorgs extends AbstractSparkAction {
 			log.debug("Number of non-Openorgs relations collected: {}", simRels.count());
 		}

-		spark
-			.createDataset(simRels.rdd(), Encoders.bean(Relation.class))
-			.write()
-			.mode(SaveMode.Overwrite)
-			.json(outputPath);
-
+		save(spark.createDataset(simRels.rdd(), Encoders.bean(Relation.class)), outputPath, SaveMode.Overwrite);
 	}

 }
--- a/dhp-workflows/dhp-dedup-openaire/src/main/java/eu/dnetlib/dhp/oa/dedup/SparkCreateMergeRels.java
+++ b/dhp-workflows/dhp-dedup-openaire/src/main/java/eu/dnetlib/dhp/oa/dedup/SparkCreateMergeRels.java
@ -155,7 +155,7 @@ public class SparkCreateMergeRels extends AbstractSparkAction {
 					(FlatMapFunction<ConnectedComponent, Relation>) cc -> ccToMergeRel(cc, dedupConf),
 					Encoders.bean(Relation.class));

-			mergeRels.write().mode(SaveMode.Overwrite).parquet(mergeRelPath);
+			saveParquet(mergeRels, mergeRelPath, SaveMode.Overwrite);

 		}
 	}
--- a/dhp-workflows/dhp-dedup-openaire/src/main/java/eu/dnetlib/dhp/oa/dedup/SparkCreateOrgsDedupRecord.java
+++ b/dhp-workflows/dhp-dedup-openaire/src/main/java/eu/dnetlib/dhp/oa/dedup/SparkCreateOrgsDedupRecord.java
@ -72,11 +72,7 @@ public class SparkCreateOrgsDedupRecord extends AbstractSparkAction {

 		final String mergeRelsPath = DedupUtility.createMergeRelPath(workingPath, actionSetId, "organization");

-		rootOrganization(spark, entityPath, mergeRelsPath)
-			.write()
-			.mode(SaveMode.Overwrite)
-			.option("compression", "gzip")
-			.json(outputPath);
+		save(rootOrganization(spark, entityPath, mergeRelsPath), outputPath, SaveMode.Overwrite);

 	}

--- a/dhp-workflows/dhp-dedup-openaire/src/main/java/eu/dnetlib/dhp/oa/dedup/SparkCreateSimRels.java
+++ b/dhp-workflows/dhp-dedup-openaire/src/main/java/eu/dnetlib/dhp/oa/dedup/SparkCreateSimRels.java
@ -82,8 +82,6 @@ public class SparkCreateSimRels extends AbstractSparkAction {
 			final String outputPath = DedupUtility.createSimRelPath(workingPath, actionSetId, subEntity);
 			removeOutputDir(spark, outputPath);

-			JavaSparkContext sc = JavaSparkContext.fromSparkContext(spark.sparkContext());
-
 			SparkDeduper deduper = new SparkDeduper(dedupConf);

 			Dataset<?> simRels = spark
--- a/dhp-workflows/dhp-dedup-openaire/src/main/java/eu/dnetlib/dhp/oa/dedup/SparkPropagateRelation.java
+++ b/dhp-workflows/dhp-dedup-openaire/src/main/java/eu/dnetlib/dhp/oa/dedup/SparkPropagateRelation.java
@ -3,23 +3,19 @@ package eu.dnetlib.dhp.oa.dedup;

 import static org.apache.spark.sql.functions.col;

-import java.util.Arrays;
-import java.util.Collections;
-import java.util.Iterator;
-import java.util.Objects;
-
-import org.apache.commons.beanutils.BeanUtils;
 import org.apache.commons.io.IOUtils;
-import org.apache.commons.lang3.StringUtils;
 import org.apache.spark.SparkConf;
-import org.apache.spark.api.java.function.FilterFunction;
 import org.apache.spark.api.java.function.MapFunction;
 import org.apache.spark.api.java.function.ReduceFunction;
 import org.apache.spark.sql.*;
+import org.apache.spark.sql.catalyst.encoders.RowEncoder;
+import org.apache.spark.sql.types.StructType;
 import org.slf4j.Logger;
 import org.slf4j.LoggerFactory;

 import eu.dnetlib.dhp.application.ArgumentApplicationParser;
+import eu.dnetlib.dhp.common.HdfsSupport;
+import eu.dnetlib.dhp.schema.common.EntityType;
 import eu.dnetlib.dhp.schema.common.ModelConstants;
 import eu.dnetlib.dhp.schema.common.ModelSupport;
 import eu.dnetlib.dhp.schema.oaf.DataInfo;
@ -70,73 +66,63 @@ public class SparkPropagateRelation extends AbstractSparkAction {
 		log.info("workingPath: '{}'", workingPath);
 		log.info("graphOutputPath: '{}'", graphOutputPath);

-		final String outputRelationPath = DedupUtility.createEntityPath(graphOutputPath, "relation");
-		removeOutputDir(spark, outputRelationPath);
-
 		Dataset<Relation> mergeRels = spark
 			.read()
 			.load(DedupUtility.createMergeRelPath(workingPath, "*", "*"))
 			.as(REL_BEAN_ENC);

 		// <mergedObjectID, dedupID>
-		Dataset<Row> mergedIds = mergeRels
+		Dataset<Row> idsToMerge = mergeRels
 			.where(col("relClass").equalTo(ModelConstants.MERGES))
 			.select(col("source").as("dedupID"), col("target").as("mergedObjectID"))
-			.distinct()
-			.cache();
+			.distinct();

 		Dataset<Row> allRels = spark
 			.read()
 			.schema(REL_BEAN_ENC.schema())
-			.json(DedupUtility.createEntityPath(graphBasePath, "relation"));
+			.json(graphBasePath + "/relation");

 		Dataset<Relation> dedupedRels = allRels
-			.joinWith(mergedIds, allRels.col("source").equalTo(mergedIds.col("mergedObjectID")), "left_outer")
-			.joinWith(mergedIds, col("_1.target").equalTo(mergedIds.col("mergedObjectID")), "left_outer")
+			.joinWith(idsToMerge, allRels.col("source").equalTo(idsToMerge.col("mergedObjectID")), "left_outer")
+			.joinWith(idsToMerge, col("_1.target").equalTo(idsToMerge.col("mergedObjectID")), "left_outer")
 			.select("_1._1", "_1._2.dedupID", "_2.dedupID")
 			.as(Encoders.tuple(REL_BEAN_ENC, Encoders.STRING(), Encoders.STRING()))
-			.flatMap(SparkPropagateRelation::addInferredRelations, REL_KRYO_ENC);
+			.map((MapFunction<Tuple3<Relation, String, String>, Relation>) t -> {
+				Relation rel = t._1();
+				String newSource = t._2();
+				String newTarget = t._3();

-		Dataset<Relation> processedRelations = distinctRelations(
-			dedupedRels.union(mergeRels.map((MapFunction<Relation, Relation>) r -> r, REL_KRYO_ENC)))
-				.filter((FilterFunction<Relation>) r -> !Objects.equals(r.getSource(), r.getTarget()));
+				if (rel.getDataInfo() == null) {
+					rel.setDataInfo(new DataInfo());
+				}

-		save(processedRelations, outputRelationPath, SaveMode.Overwrite);
-	}
+				if (newSource != null || newTarget != null) {
+					rel.getDataInfo().setDeletedbyinference(false);

-	private static Iterator<Relation> addInferredRelations(Tuple3<Relation, String, String> t) throws Exception {
-		Relation existingRel = t._1();
-		String newSource = t._2();
-		String newTarget = t._3();
+					if (newSource != null)
+						rel.setSource(newSource);

-		if (newSource == null && newTarget == null) {
-			return Collections.singleton(t._1()).iterator();
-		}
+					if (newTarget != null)
+						rel.setTarget(newTarget);
+				}

-		// update existing relation
-		if (existingRel.getDataInfo() == null) {
-			existingRel.setDataInfo(new DataInfo());
-		}
-		existingRel.getDataInfo().setDeletedbyinference(true);
+				return rel;
+			}, REL_BEAN_ENC);

-		// Create new relation inferred by dedupIDs
-		Relation inferredRel = (Relation) BeanUtils.cloneBean(existingRel);
+		// ids of records that are both not deletedbyinference and not invisible
+		Dataset<Row> ids = validIds(spark, graphBasePath);

-		inferredRel.setDataInfo((DataInfo) BeanUtils.cloneBean(existingRel.getDataInfo()));
-		inferredRel.getDataInfo().setDeletedbyinference(false);
+		// filter relations that point to valid records, can force them to be visible
+		Dataset<Relation> cleanedRels = dedupedRels
+			.join(ids, col("source").equalTo(ids.col("id")), "leftsemi")
+			.join(ids, col("target").equalTo(ids.col("id")), "leftsemi")
+			.as(REL_BEAN_ENC)
+			.map((MapFunction<Relation, Relation>) r -> {
+				r.getDataInfo().setInvisible(false);
+				return r;
+			}, REL_KRYO_ENC);

-		if (newSource != null)
-			inferredRel.setSource(newSource);
-
-		if (newTarget != null)
-			inferredRel.setTarget(newTarget);
-
-		return Arrays.asList(existingRel, inferredRel).iterator();
-	}
-
-	private Dataset<Relation> distinctRelations(Dataset<Relation> rels) {
-		return rels
-			.filter(getRelationFilterFunction())
+		Dataset<Relation> distinctRels = cleanedRels
 			.groupByKey(
 				(MapFunction<Relation, String>) r -> String
 					.join(" ", r.getSource(), r.getTarget(), r.getRelType(), r.getSubRelType(), r.getRelClass()),
@ -146,13 +132,33 @@ public class SparkPropagateRelation extends AbstractSparkAction {
 				return b;
 			})
 			.map((MapFunction<Tuple2<String, Relation>, Relation>) Tuple2::_2, REL_BEAN_ENC);
+
+		final String outputRelationPath = graphOutputPath + "/relation";
+		removeOutputDir(spark, outputRelationPath);
+		save(
+			distinctRels
+				.union(mergeRels)
+				.filter("source != target AND dataInfo.deletedbyinference != true AND dataInfo.invisible != true"),
+			outputRelationPath,
+			SaveMode.Overwrite);
 	}

-	private FilterFunction<Relation> getRelationFilterFunction() {
-		return r -> StringUtils.isNotBlank(r.getSource()) ||
-			StringUtils.isNotBlank(r.getTarget()) ||
-			StringUtils.isNotBlank(r.getRelType()) ||
-			StringUtils.isNotBlank(r.getSubRelType()) ||
-			StringUtils.isNotBlank(r.getRelClass());
+	static Dataset<Row> validIds(SparkSession spark, String graphBasePath) {
+		StructType idsSchema = StructType
+			.fromDDL("`id` STRING, `dataInfo` STRUCT<`deletedbyinference`:BOOLEAN,`invisible`:BOOLEAN>");
+
+		Dataset<Row> allIds = spark.emptyDataset(RowEncoder.apply(idsSchema));
+
+		for (EntityType entityType : ModelSupport.entityTypes.keySet()) {
+			String entityPath = graphBasePath + '/' + entityType.name();
+			if (HdfsSupport.exists(entityPath, spark.sparkContext().hadoopConfiguration())) {
+				allIds = allIds.union(spark.read().schema(idsSchema).json(entityPath));
+			}
+		}
+
+		return allIds
+			.filter("dataInfo.deletedbyinference != true AND dataInfo.invisible != true")
+			.select("id")
+			.distinct();
 	}
 }
--- a/dhp-workflows/dhp-dedup-openaire/src/main/java/eu/dnetlib/dhp/oa/dedup/SparkWhitelistSimRels.java
+++ b/dhp-workflows/dhp-dedup-openaire/src/main/java/eu/dnetlib/dhp/oa/dedup/SparkWhitelistSimRels.java
@ -67,8 +67,6 @@ public class SparkWhitelistSimRels extends AbstractSparkAction {
 		log.info("workingPath:   '{}'", workingPath);
 		log.info("whiteListPath: '{}'", whiteListPath);

-		JavaSparkContext sc = JavaSparkContext.fromSparkContext(spark.sparkContext());
-
 		// file format: source####target
 		Dataset<Row> whiteListRels = spark
 			.read()
--- a/dhp-workflows/dhp-dedup-openaire/src/main/resources/eu/dnetlib/dhp/oa/dedup/cleanRelation_parameters.json
+++ b/dhp-workflows/dhp-dedup-openaire/src/main/resources/eu/dnetlib/dhp/oa/dedup/cleanRelation_parameters.json
@ -1,20 +0,0 @@
-[
-  {
-    "paramName": "i",
-    "paramLongName": "graphBasePath",
-    "paramDescription": "the base path of raw graph",
-    "paramRequired": true
-  },
-  {
-    "paramName": "w",
-    "paramLongName": "inputPath",
-    "paramDescription": "the path to the input relation to cleanup",
-    "paramRequired": true
-  },
-  {
-    "paramName": "o",
-    "paramLongName": "outputPath",
-    "paramDescription": "the path of the output relation cleaned",
-    "paramRequired": true
-  }
-]
--- a/dhp-workflows/dhp-dedup-openaire/src/main/resources/eu/dnetlib/dhp/oa/dedup/consistency/oozie_app/workflow.xml
+++ b/dhp-workflows/dhp-dedup-openaire/src/main/resources/eu/dnetlib/dhp/oa/dedup/consistency/oozie_app/workflow.xml
@ -100,35 +100,9 @@
                --conf spark.sql.shuffle.partitions=15000
            </spark-opts>
            <arg>--graphBasePath</arg><arg>${graphBasePath}</arg>
-            <arg>--graphOutputPath</arg><arg>${workingPath}/propagaterelation/</arg>
+            <arg>--graphOutputPath</arg><arg>${graphOutputPath}</arg>
            <arg>--workingPath</arg><arg>${workingPath}</arg>
        </spark>
-        <ok to="CleanRelation"/>
-        <error to="Kill"/>
-    </action>
-
-    <action name="CleanRelation">
-        <spark xmlns="uri:oozie:spark-action:0.2">
-            <master>yarn</master>
-            <mode>cluster</mode>
-            <name>Clean Relations</name>
-            <class>eu.dnetlib.dhp.oa.dedup.SparkCleanRelation</class>
-            <jar>dhp-dedup-openaire-${projectVersion}.jar</jar>
-            <spark-opts>
-                --executor-memory=${sparkExecutorMemory}
-                --conf spark.executor.memoryOverhead=${sparkExecutorMemoryOverhead}
-                --executor-cores=${sparkExecutorCores}
-                --driver-memory=${sparkDriverMemory}
-                --conf spark.extraListeners=${spark2ExtraListeners}
-                --conf spark.sql.queryExecutionListeners=${spark2SqlQueryExecutionListeners}
-                --conf spark.yarn.historyServer.address=${spark2YarnHistoryServerAddress}
-                --conf spark.eventLog.dir=${nameNode}${spark2EventLogDir}
-                --conf spark.sql.shuffle.partitions=15000
-            </spark-opts>
-            <arg>--graphBasePath</arg><arg>${graphBasePath}</arg>
-            <arg>--inputPath</arg><arg>${workingPath}/propagaterelation/relation</arg>
-            <arg>--outputPath</arg><arg>${graphOutputPath}/relation</arg>
-        </spark>
        <ok to="group_entities"/>
        <error to="Kill"/>
    </action>
@ -152,31 +126,7 @@
                --conf spark.sql.shuffle.partitions=15000
            </spark-opts>
            <arg>--graphInputPath</arg><arg>${graphBasePath}</arg>
-            <arg>--outputPath</arg><arg>${workingPath}/grouped_entities</arg>
-        </spark>
-        <ok to="dispatch_entities"/>
-        <error to="Kill"/>
-    </action>
-
-    <action name="dispatch_entities">
-        <spark xmlns="uri:oozie:spark-action:0.2">
-            <master>yarn</master>
-            <mode>cluster</mode>
-            <name>Dispatch grouped entitities</name>
-            <class>eu.dnetlib.dhp.oa.merge.DispatchEntitiesSparkJob</class>
-            <jar>dhp-dedup-openaire-${projectVersion}.jar</jar>
-            <spark-opts>
-                --executor-memory=${sparkExecutorMemory}
-                --conf spark.executor.memoryOverhead=${sparkExecutorMemoryOverhead}
-                --executor-cores=${sparkExecutorCores}
-                --driver-memory=${sparkDriverMemory}
-                --conf spark.extraListeners=${spark2ExtraListeners}
-                --conf spark.sql.queryExecutionListeners=${spark2SqlQueryExecutionListeners}
-                --conf spark.yarn.historyServer.address=${spark2YarnHistoryServerAddress}
-                --conf spark.eventLog.dir=${nameNode}${spark2EventLogDir}
-                --conf spark.sql.shuffle.partitions=7680
-            </spark-opts>
-            <arg>--inputPath</arg><arg>${workingPath}/grouped_entities</arg>
+            <arg>--checkpointPath</arg><arg>${workingPath}/grouped_entities</arg>
            <arg>--outputPath</arg><arg>${graphOutputPath}</arg>
            <arg>--filterInvisible</arg><arg>${filterInvisible}</arg>
        </spark>
--- a/dhp-workflows/dhp-dedup-openaire/src/test/java/eu/dnetlib/dhp/oa/dedup/SparkDedupTest.java
+++ b/dhp-workflows/dhp-dedup-openaire/src/test/java/eu/dnetlib/dhp/oa/dedup/SparkDedupTest.java
@ -3,7 +3,6 @@ package eu.dnetlib.dhp.oa.dedup;

 import static java.nio.file.Files.createTempDirectory;

-import static org.apache.spark.sql.functions.col;
 import static org.apache.spark.sql.functions.count;
 import static org.junit.jupiter.api.Assertions.*;
 import static org.mockito.Mockito.lenient;
@ -23,14 +22,13 @@ import java.util.stream.Collectors;
 import org.apache.commons.io.FileUtils;
 import org.apache.commons.io.IOUtils;
 import org.apache.spark.SparkConf;
-import org.apache.spark.api.java.JavaPairRDD;
-import org.apache.spark.api.java.JavaRDD;
 import org.apache.spark.api.java.JavaSparkContext;
 import org.apache.spark.api.java.function.FilterFunction;
 import org.apache.spark.api.java.function.MapFunction;
-import org.apache.spark.api.java.function.PairFunction;
-import org.apache.spark.sql.*;
 import org.apache.spark.sql.Dataset;
+import org.apache.spark.sql.Encoders;
+import org.apache.spark.sql.Row;
+import org.apache.spark.sql.SparkSession;
 import org.junit.jupiter.api.*;
 import org.junit.jupiter.api.extension.ExtendWith;
 import org.mockito.Mock;
@ -46,8 +44,6 @@ import eu.dnetlib.dhp.schema.common.ModelConstants;
 import eu.dnetlib.dhp.schema.oaf.*;
 import eu.dnetlib.enabling.is.lookup.rmi.ISLookUpException;
 import eu.dnetlib.enabling.is.lookup.rmi.ISLookUpService;
-import eu.dnetlib.pace.util.MapDocumentUtil;
-import scala.Tuple2;

@ExtendWith(MockitoExtension.class)
@TestMethodOrder(MethodOrderer.OrderAnnotation.class)
@ -62,6 +58,8 @@ public class SparkDedupTest implements Serializable {
 	private static String testGraphBasePath;
 	private static String testOutputBasePath;
 	private static String testDedupGraphBasePath;
+	private static String testConsistencyGraphBasePath;
+
 	private static final String testActionSetId = "test-orchestrator";
 	private static String whitelistPath;
 	private static List<String> whiteList;
@ -75,6 +73,7 @@ public class SparkDedupTest implements Serializable {
 			.get(SparkDedupTest.class.getResource("/eu/dnetlib/dhp/dedup/entities").toURI())
 			.toFile()
 			.getAbsolutePath();
+
 		testOutputBasePath = createTempDirectory(SparkDedupTest.class.getSimpleName() + "-")
 			.toAbsolutePath()
 			.toString();
@ -83,6 +82,10 @@ public class SparkDedupTest implements Serializable {
 			.toAbsolutePath()
 			.toString();

+		testConsistencyGraphBasePath = createTempDirectory(SparkDedupTest.class.getSimpleName() + "-")
+			.toAbsolutePath()
+			.toString();
+
 		whitelistPath = Paths
 			.get(SparkDedupTest.class.getResource("/eu/dnetlib/dhp/dedup/whitelist.simrels.txt").toURI())
 			.toFile()
@ -674,22 +677,45 @@ public class SparkDedupTest implements Serializable {
 		assertEquals(mergedOrp, deletedOrp);
 	}

+	@Test
+	@Order(6)
+	void copyRelationsNoOpenorgsTest() throws Exception {
+
+		ArgumentApplicationParser parser = new ArgumentApplicationParser(
+			IOUtils
+				.toString(
+					SparkCopyRelationsNoOpenorgs.class
+						.getResourceAsStream(
+							"/eu/dnetlib/dhp/oa/dedup/updateEntity_parameters.json")));
+		parser
+			.parseArgument(
+				new String[] {
+					"-i", testGraphBasePath, "-w", testOutputBasePath, "-o", testDedupGraphBasePath
+				});
+
+		new SparkCopyRelationsNoOpenorgs(parser, spark).run(isLookUpService);
+
+		final Dataset<Row> outputRels = spark.read().text(testDedupGraphBasePath + "/relation");
+
+		System.out.println(outputRels.count());
+		// assertEquals(2382, outputRels.count());
+	}
+
 	@Test
 	@Order(7)
 	void propagateRelationTest() throws Exception {

 		ArgumentApplicationParser parser = new ArgumentApplicationParser(
 			classPathResourceAsString("/eu/dnetlib/dhp/oa/dedup/propagateRelation_parameters.json"));
-		String outputRelPath = testDedupGraphBasePath + "/propagaterelation";
 		parser
 			.parseArgument(
 				new String[] {
-					"-i", testGraphBasePath, "-w", testOutputBasePath, "-o", outputRelPath
+					"-i", testDedupGraphBasePath, "-w", testOutputBasePath, "-o", testConsistencyGraphBasePath
 				});

 		new SparkPropagateRelation(parser, spark).run(isLookUpService);

-		long relations = jsc.textFile(outputRelPath + "/relation").count();
+		long relations = jsc.textFile(testDedupGraphBasePath + "/relation").count();

 //		assertEquals(4860, relations);
 		System.out.println("relations = " + relations);
@ -699,95 +725,52 @@ public class SparkDedupTest implements Serializable {
 			.read()
 			.load(DedupUtility.createMergeRelPath(testOutputBasePath, "*", "*"))
 			.as(Encoders.bean(Relation.class));
-		final JavaPairRDD<String, String> mergedIds = mergeRels
-			.where("relClass == 'merges'")
-			.select(mergeRels.col("target"))
-			.distinct()
-			.toJavaRDD()
-			.mapToPair(
-				(PairFunction<Row, String, String>) r -> new Tuple2<String, String>(r.getString(0), "d"));

-		JavaRDD<String> toCheck = jsc
-			.textFile(outputRelPath + "/relation")
-			.mapToPair(json -> new Tuple2<>(MapDocumentUtil.getJPathString("$.source", json), json))
-			.join(mergedIds)
-			.map(t -> t._2()._1())
-			.mapToPair(json -> new Tuple2<>(MapDocumentUtil.getJPathString("$.target", json), json))
-			.join(mergedIds)
-			.map(t -> t._2()._1());
+		Dataset<Row> inputRels = spark
+			.read()
+			.json(testDedupGraphBasePath + "/relation");

-		long deletedbyinference = toCheck.filter(this::isDeletedByInference).count();
-		long updated = toCheck.count();
+		Dataset<Row> outputRels = spark
+			.read()
+			.json(testConsistencyGraphBasePath + "/relation");

-		assertEquals(updated, deletedbyinference);
+		assertEquals(
+			0, outputRels
+				.filter("dataInfo.deletedbyinference == true OR dataInfo.invisible == true")
+				.count());
+
+		assertEquals(
+			5, outputRels
+				.filter("relClass NOT IN ('merges', 'isMergedIn')")
+				.count());
+
+		assertEquals(5 + mergeRels.count(), outputRels.count());
 	}

 	@Test
 	@Order(8)
-	void testCleanBaseRelations() throws Exception {
-		ArgumentApplicationParser parser = new ArgumentApplicationParser(
-			classPathResourceAsString("/eu/dnetlib/dhp/oa/dedup/cleanRelation_parameters.json"));
-
-		// append dangling relations to be cleaned up
+	void testCleanedPropagatedRelations() throws Exception {
 		Dataset<Row> df_before = spark
 			.read()
 			.schema(Encoders.bean(Relation.class).schema())
-			.json(testGraphBasePath + "/relation");
-		Dataset<Row> df_input = df_before
-			.unionByName(df_before.drop("source").withColumn("source", functions.lit("n/a")))
-			.unionByName(df_before.drop("target").withColumn("target", functions.lit("n/a")));
-		df_input.write().mode(SaveMode.Overwrite).json(testOutputBasePath + "_tmp");
-
-		parser
-			.parseArgument(
-				new String[] {
-					"--graphBasePath", testGraphBasePath,
-					"--inputPath", testGraphBasePath + "/relation",
-					"--outputPath", testDedupGraphBasePath + "/relation"
-				});
-
-		new SparkCleanRelation(parser, spark).run(isLookUpService);
+			.json(testDedupGraphBasePath + "/relation");

 		Dataset<Row> df_after = spark
 			.read()
 			.schema(Encoders.bean(Relation.class).schema())
-			.json(testDedupGraphBasePath + "/relation");
-
-		assertNotEquals(df_before.count(), df_input.count());
-		assertNotEquals(df_input.count(), df_after.count());
-		assertEquals(5, df_after.count());
-	}
-
-	@Test
-	@Order(9)
-	void testCleanDedupedRelations() throws Exception {
-		ArgumentApplicationParser parser = new ArgumentApplicationParser(
-			classPathResourceAsString("/eu/dnetlib/dhp/oa/dedup/cleanRelation_parameters.json"));
-
-		String inputRelPath = testDedupGraphBasePath + "/propagaterelation/relation";
-
-		// append dangling relations to be cleaned up
-		Dataset<Row> df_before = spark.read().schema(Encoders.bean(Relation.class).schema()).json(inputRelPath);
-
-		df_before.filter(col("dataInfo.deletedbyinference").notEqual(true)).show(50, false);
-
-		parser
-			.parseArgument(
-				new String[] {
-					"--graphBasePath", testGraphBasePath,
-					"--inputPath", inputRelPath,
-					"--outputPath", testDedupGraphBasePath + "/relation"
-				});
-
-		new SparkCleanRelation(parser, spark).run(isLookUpService);
-
-		Dataset<Row> df_after = spark
-			.read()
-			.schema(Encoders.bean(Relation.class).schema())
-			.json(testDedupGraphBasePath + "/relation");
+			.json(testConsistencyGraphBasePath + "/relation");

 		assertNotEquals(df_before.count(), df_after.count());
-		assertEquals(0, df_after.count());
+
+		assertEquals(
+			0, df_after
+				.filter("dataInfo.deletedbyinference == true OR dataInfo.invisible == true")
+				.count());
+
+		assertEquals(
+			5, df_after
+				.filter("relClass NOT IN ('merges', 'isMergedIn')")
+				.count());
 	}

 	@Test
@ -813,6 +796,7 @@ public class SparkDedupTest implements Serializable {
 	public static void finalCleanUp() throws IOException {
 		FileUtils.deleteDirectory(new File(testOutputBasePath));
 		FileUtils.deleteDirectory(new File(testDedupGraphBasePath));
+		FileUtils.deleteDirectory(new File(testConsistencyGraphBasePath));
 	}

 	public boolean isDeletedByInference(String s) {
--- a/dhp-workflows/dhp-dedup-openaire/src/test/java/eu/dnetlib/dhp/oa/dedup/SparkOpenorgsProvisionTest.java
+++ b/dhp-workflows/dhp-dedup-openaire/src/test/java/eu/dnetlib/dhp/oa/dedup/SparkOpenorgsProvisionTest.java
@ -3,6 +3,7 @@ package eu.dnetlib.dhp.oa.dedup;

 import static java.nio.file.Files.createTempDirectory;

+import static org.apache.spark.sql.functions.col;
 import static org.junit.jupiter.api.Assertions.assertEquals;
 import static org.mockito.Mockito.lenient;

@ -15,10 +16,6 @@ import java.nio.file.Paths;
 import org.apache.commons.io.FileUtils;
 import org.apache.commons.io.IOUtils;
 import org.apache.spark.SparkConf;
-import org.apache.spark.api.java.JavaPairRDD;
-import org.apache.spark.api.java.JavaRDD;
-import org.apache.spark.api.java.JavaSparkContext;
-import org.apache.spark.api.java.function.PairFunction;
 import org.apache.spark.sql.Dataset;
 import org.apache.spark.sql.Encoders;
 import org.apache.spark.sql.Row;
@ -33,8 +30,6 @@ import eu.dnetlib.dhp.application.ArgumentApplicationParser;
 import eu.dnetlib.dhp.schema.oaf.Relation;
 import eu.dnetlib.enabling.is.lookup.rmi.ISLookUpException;
 import eu.dnetlib.enabling.is.lookup.rmi.ISLookUpService;
-import eu.dnetlib.pace.util.MapDocumentUtil;
-import scala.Tuple2;

@ExtendWith(MockitoExtension.class)
@TestMethodOrder(MethodOrderer.OrderAnnotation.class)
@ -44,11 +39,11 @@ public class SparkOpenorgsProvisionTest implements Serializable {
 	ISLookUpService isLookUpService;

 	private static SparkSession spark;
-	private static JavaSparkContext jsc;

 	private static String testGraphBasePath;
 	private static String testOutputBasePath;
 	private static String testDedupGraphBasePath;
+	private static String testConsistencyGraphBasePath;
 	private static final String testActionSetId = "test-orchestrator";

 	@BeforeAll
@ -64,6 +59,9 @@ public class SparkOpenorgsProvisionTest implements Serializable {
 		testDedupGraphBasePath = createTempDirectory(SparkOpenorgsProvisionTest.class.getSimpleName() + "-")
 			.toAbsolutePath()
 			.toString();
+		testConsistencyGraphBasePath = createTempDirectory(SparkOpenorgsProvisionTest.class.getSimpleName() + "-")
+			.toAbsolutePath()
+			.toString();

 		FileUtils.deleteDirectory(new File(testOutputBasePath));
 		FileUtils.deleteDirectory(new File(testDedupGraphBasePath));
@ -76,8 +74,13 @@ public class SparkOpenorgsProvisionTest implements Serializable {
 			.master("local[*]")
 			.config(conf)
 			.getOrCreate();
+	}

-		jsc = JavaSparkContext.fromSparkContext(spark.sparkContext());
+	@AfterAll
+	public static void finalCleanUp() throws IOException {
+		FileUtils.deleteDirectory(new File(testOutputBasePath));
+		FileUtils.deleteDirectory(new File(testDedupGraphBasePath));
+		FileUtils.deleteDirectory(new File(testConsistencyGraphBasePath));
 	}

 	@BeforeEach
@ -186,26 +189,21 @@ public class SparkOpenorgsProvisionTest implements Serializable {

 		new SparkUpdateEntity(parser, spark).run(isLookUpService);

-		long organizations = jsc.textFile(testDedupGraphBasePath + "/organization").count();
+		Dataset<Row> organizations = spark.read().json(testDedupGraphBasePath + "/organization");

-		long mergedOrgs = spark
+		Dataset<Row> mergedOrgs = spark
 			.read()
 			.load(testOutputBasePath + "/" + testActionSetId + "/organization_mergerel")
-			.as(Encoders.bean(Relation.class))
 			.where("relClass=='merges'")
-			.javaRDD()
-			.map(Relation::getTarget)
-			.distinct()
-			.count();
+			.select("target")
+			.distinct();

-		assertEquals(80, organizations);
+		assertEquals(80, organizations.count());

-		long deletedOrgs = jsc
-			.textFile(testDedupGraphBasePath + "/organization")
-			.filter(this::isDeletedByInference)
-			.count();
+		Dataset<Row> deletedOrgs = organizations
+			.filter("dataInfo.deletedbyinference = TRUE");

-		assertEquals(mergedOrgs, deletedOrgs);
+		assertEquals(mergedOrgs.count(), deletedOrgs.count());
 	}

 	@Test
@ -226,10 +224,9 @@ public class SparkOpenorgsProvisionTest implements Serializable {

 		new SparkCopyRelationsNoOpenorgs(parser, spark).run(isLookUpService);

-		final JavaRDD<String> rels = jsc.textFile(testDedupGraphBasePath + "/relation");
-
-		assertEquals(2382, rels.count());
+		final Dataset<Row> outputRels = spark.read().text(testDedupGraphBasePath + "/relation");

+		assertEquals(2382, outputRels.count());
 	}

 	@Test
@ -244,51 +241,41 @@ public class SparkOpenorgsProvisionTest implements Serializable {
 		parser
 			.parseArgument(
 				new String[] {
-					"-i", testGraphBasePath, "-w", testOutputBasePath, "-o", testDedupGraphBasePath
+					"-i", testDedupGraphBasePath, "-w", testOutputBasePath, "-o", testConsistencyGraphBasePath
 				});

 		new SparkPropagateRelation(parser, spark).run(isLookUpService);

-		long relations = jsc.textFile(testDedupGraphBasePath + "/relation").count();
-
-		assertEquals(4896, relations);
-
-		// check deletedbyinference
 		final Dataset<Relation> mergeRels = spark
 			.read()
 			.load(DedupUtility.createMergeRelPath(testOutputBasePath, "*", "*"))
 			.as(Encoders.bean(Relation.class));
-		final JavaPairRDD<String, String> mergedIds = mergeRels
+
+		Dataset<Row> inputRels = spark
+			.read()
+			.json(testDedupGraphBasePath + "/relation");
+
+		Dataset<Row> outputRels = spark
+			.read()
+			.json(testConsistencyGraphBasePath + "/relation");
+
+		final Dataset<Row> mergedIds = mergeRels
 			.where("relClass == 'merges'")
-			.select(mergeRels.col("target"))
-			.distinct()
-			.toJavaRDD()
-			.mapToPair(
-				(PairFunction<Row, String, String>) r -> new Tuple2<String, String>(r.getString(0), "d"));
+			.select(col("target").as("id"))
+			.distinct();

-		JavaRDD<String> toCheck = jsc
-			.textFile(testDedupGraphBasePath + "/relation")
-			.mapToPair(json -> new Tuple2<>(MapDocumentUtil.getJPathString("$.source", json), json))
-			.join(mergedIds)
-			.map(t -> t._2()._1())
-			.mapToPair(json -> new Tuple2<>(MapDocumentUtil.getJPathString("$.target", json), json))
-			.join(mergedIds)
-			.map(t -> t._2()._1());
+		Dataset<Row> toUpdateRels = inputRels
+			.as("rel")
+			.join(mergedIds.as("s"), col("rel.source").equalTo(col("s.id")), "left_outer")
+			.join(mergedIds.as("t"), col("rel.target").equalTo(col("t.id")), "left_outer")
+			.filter("s.id IS NOT NULL OR t.id IS NOT NULL")
+			.distinct();

-		long deletedbyinference = toCheck.filter(this::isDeletedByInference).count();
-		long updated = toCheck.count();
+		Dataset<Row> updatedRels = inputRels
+			.select("source", "target", "relClass")
+			.except(outputRels.select("source", "target", "relClass"));

-		assertEquals(updated, deletedbyinference);
+		assertEquals(toUpdateRels.count(), updatedRels.count());
+		assertEquals(140, outputRels.count());
 	}
-
-	@AfterAll
-	public static void finalCleanUp() throws IOException {
-		FileUtils.deleteDirectory(new File(testOutputBasePath));
-		FileUtils.deleteDirectory(new File(testDedupGraphBasePath));
-	}
-
-	public boolean isDeletedByInference(String s) {
-		return s.contains("\"deletedbyinference\":true");
-	}
-
 }
--- a/dhp-workflows/dhp-distcp/pom.xml
+++ b/dhp-workflows/dhp-distcp/pom.xml
@ -1,13 +0,0 @@
-<?xml version="1.0" encoding="UTF-8"?>
-<project xmlns="http://maven.apache.org/POM/4.0.0" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://maven.apache.org/POM/4.0.0 http://maven.apache.org/xsd/maven-4.0.0.xsd">
-    <parent>
-        <artifactId>dhp-workflows</artifactId>
-        <groupId>eu.dnetlib.dhp</groupId>
-        <version>1.2.5-SNAPSHOT</version>
-    </parent>
-    <modelVersion>4.0.0</modelVersion>
-
-    <artifactId>dhp-distcp</artifactId>
-
-
-</project>
--- a/dhp-workflows/dhp-distcp/src/main/resources/eu/dnetlib/dhp/distcp/oozie_app/config-default.xml
+++ b/dhp-workflows/dhp-distcp/src/main/resources/eu/dnetlib/dhp/distcp/oozie_app/config-default.xml
@ -1,18 +0,0 @@
-<configuration>
-    <property>
-        <name>jobTracker</name>
-        <value>yarnRM</value>
-    </property>
-    <property>
-        <name>nameNode</name>
-        <value>hdfs://nameservice1</value>
-    </property>
-    <property>
-        <name>sourceNN</name>
-        <value>webhdfs://namenode2.hadoop.dm.openaire.eu:50071</value>
-    </property>
-    <property>
-        <name>oozie.use.system.libpath</name>
-        <value>true</value>
-    </property>
-</configuration>
--- a/dhp-workflows/dhp-distcp/src/main/resources/eu/dnetlib/dhp/distcp/oozie_app/workflow.xml
+++ b/dhp-workflows/dhp-distcp/src/main/resources/eu/dnetlib/dhp/distcp/oozie_app/workflow.xml
@ -1,46 +0,0 @@
-<workflow-app name="distcp" xmlns="uri:oozie:workflow:0.5">
-    <parameters>
-        <property>
-            <name>sourceNN</name>
-            <description>the source name node</description>
-        </property>
-        <property>
-            <name>sourcePath</name>
-            <description>the source path</description>
-        </property>
-        <property>
-            <name>targetPath</name>
-            <description>the target path</description>
-        </property>
-        <property>
-            <name>hbase_dump_distcp_memory_mb</name>
-            <value>6144</value>
-            <description>memory for distcp action copying InfoSpace dump from remote cluster</description>
-        </property>
-        <property>
-            <name>hbase_dump_distcp_num_maps</name>
-            <value>1</value>
-            <description>maximum number of simultaneous copies of InfoSpace dump from remote location</description>
-        </property>
-    </parameters>
-
-    <start to="distcp"/>
-
-    <kill name="Kill">
-        <message>Action failed, error message[${wf:errorMessage(wf:lastErrorNode())}]</message>
-    </kill>
-
-    <action name="distcp">
-        <distcp xmlns="uri:oozie:distcp-action:0.2">
-            <arg>-Dmapreduce.map.memory.mb=${hbase_dump_distcp_memory_mb}</arg>
-            <arg>-pb</arg>
-            <arg>-m ${hbase_dump_distcp_num_maps}</arg>
-            <arg>${sourceNN}/${sourcePath}</arg>
-            <arg>${nameNode}/${targetPath}</arg>
-        </distcp>
-        <ok to="End" />
-        <error to="Kill" />
-    </action>
-
-    <end name="End"/>
-</workflow-app>
--- a/dhp-workflows/dhp-doiboost/src/main/resources/eu/dnetlib/dhp/doiboost/crossref/irish_funder.json
+++ b/dhp-workflows/dhp-doiboost/src/main/resources/eu/dnetlib/dhp/doiboost/crossref/irish_funder.json
@ -0,0 +1,940 @@
+[
+  {
+    "id": "100007630",
+    "uri": "http://dx.doi.org/10.13039/100007630",
+    "name": "College of Engineering and Informatics, National University of Ireland, Galway",
+    "synonym": []
+  },
+  {
+    "id": "100007731",
+    "uri": "http://dx.doi.org/10.13039/100007731",
+    "name": "Endo International",
+    "synonym": []
+  },
+  {
+    "id": "100008099",
+    "uri": "http://dx.doi.org/10.13039/100008099",
+    "name": "Food Safety Authority of Ireland",
+    "synonym": []
+  },
+  {
+    "id": "100008124",
+    "uri": "http://dx.doi.org/10.13039/100008124",
+    "name": "Department of Jobs, Enterprise and Innovation",
+    "synonym": []
+  },
+  {
+    "id": "100009098",
+    "uri": "http://dx.doi.org/10.13039/100009098",
+    "name": "Department of Foreign Affairs and Trade, Ireland",
+    "synonym": []
+  },
+  {
+    "id": "100009099",
+    "uri": "http://dx.doi.org/10.13039/100009099",
+    "name": "Irish Aid",
+    "synonym": []
+  },
+  {
+    "id": "100009770",
+    "uri": "http://dx.doi.org/10.13039/100009770",
+    "name": "National University of Ireland",
+    "synonym": []
+  },
+  {
+    "id": "100009985",
+    "uri": "http://dx.doi.org/10.13039/100009985",
+    "name": "Parkinson's Association of Ireland",
+    "synonym": []
+  },
+  {
+    "id": "100010399",
+    "uri": "http://dx.doi.org/10.13039/100010399",
+    "name": "European Society of Cataract and Refractive Surgeons",
+    "synonym": []
+  },
+  {
+    "id": "100010414",
+    "uri": "http://dx.doi.org/10.13039/100010414",
+    "name": "Health Research Board",
+    "synonym": [
+      "501100001590"
+    ]
+  },
+  {
+    "id": "100010546",
+    "uri": "http://dx.doi.org/10.13039/100010546",
+    "name": "Deparment of Children and Youth Affairs, Ireland",
+    "synonym": []
+  },
+  {
+    "id": "100010993",
+    "uri": "http://dx.doi.org/10.13039/100010993",
+    "name": "Irish Nephrology Society",
+    "synonym": []
+  },
+  {
+    "id": "100011062",
+    "uri": "http://dx.doi.org/10.13039/100011062",
+    "name": "Asian Spinal Cord Network",
+    "synonym": []
+  },
+  {
+    "id": "100011096",
+    "uri": "http://dx.doi.org/10.13039/100011096",
+    "name": "Jazz Pharmaceuticals",
+    "synonym": []
+  },
+  {
+    "id": "100011396",
+    "uri": "http://dx.doi.org/10.13039/100011396",
+    "name": "Irish College of General Practitioners",
+    "synonym": []
+  },
+  {
+    "id": "100012734",
+    "uri": "http://dx.doi.org/10.13039/100012734",
+    "name": "Department for Culture, Heritage and the Gaeltacht, Ireland",
+    "synonym": []
+  },
+  {
+    "id": "100012754",
+    "uri": "http://dx.doi.org/10.13039/100012754",
+    "name": "Horizon Pharma",
+    "synonym": []
+  },
+  {
+    "id": "100012891",
+    "uri": "http://dx.doi.org/10.13039/100012891",
+    "name": "Medical Research Charities Group",
+    "synonym": []
+  },
+  {
+    "id": "100012919",
+    "uri": "http://dx.doi.org/10.13039/100012919",
+    "name": "Epilepsy Ireland",
+    "synonym": []
+  },
+  {
+    "id": "100012920",
+    "uri": "http://dx.doi.org/10.13039/100012920",
+    "name": "GLEN",
+    "synonym": []
+  },
+  {
+    "id": "100012921",
+    "uri": "http://dx.doi.org/10.13039/100012921",
+    "name": "Royal College of Surgeons in Ireland",
+    "synonym": []
+  },
+  {
+    "id": "100013029",
+    "uri": "http://dx.doi.org/10.13039/100013029",
+    "name": "Iris O'Brien Foundation",
+    "synonym": []
+  },
+  {
+    "id": "100013206",
+    "uri": "http://dx.doi.org/10.13039/100013206",
+    "name": "Food Institutional Research Measure",
+    "synonym": []
+  },
+  {
+    "id": "100013381",
+    "uri": "http://dx.doi.org/10.13039/100013381",
+    "name": "Irish Phytochemical Food Network",
+    "synonym": []
+  },
+  {
+    "id": "100013433",
+    "uri": "http://dx.doi.org/10.13039/100013433",
+    "name": "Transport Infrastructure Ireland",
+    "synonym": []
+  },
+  {
+    "id": "100013461",
+    "uri": "http://dx.doi.org/10.13039/100013461",
+    "name": "Arts and Disability Ireland",
+    "synonym": []
+  },
+  {
+    "id": "100013548",
+    "uri": "http://dx.doi.org/10.13039/100013548",
+    "name": "Filmbase",
+    "synonym": []
+  },
+  {
+    "id": "100013917",
+    "uri": "http://dx.doi.org/10.13039/100013917",
+    "name": "Society for Musicology in Ireland",
+    "synonym": []
+  },
+  {
+    "id": "100014251",
+    "uri": "http://dx.doi.org/10.13039/100014251",
+    "name": "Humanities in the European Research Area",
+    "synonym": []
+  },
+  {
+    "id": "100014364",
+    "uri": "http://dx.doi.org/10.13039/100014364",
+    "name": "National Children's Research Centre",
+    "synonym": []
+  },
+  {
+    "id": "100014384",
+    "uri": "http://dx.doi.org/10.13039/100014384",
+    "name": "Amarin Corporation",
+    "synonym": []
+  },
+  {
+    "id": "100014902",
+    "uri": "http://dx.doi.org/10.13039/100014902",
+    "name": "Irish Association for Cancer Research",
+    "synonym": []
+  },
+  {
+    "id": "100015023",
+    "uri": "http://dx.doi.org/10.13039/100015023",
+    "name": "Ireland Funds",
+    "synonym": []
+  },
+  {
+    "id": "100015037",
+    "uri": "http://dx.doi.org/10.13039/100015037",
+    "name": "Simon Cumbers Media Fund",
+    "synonym": []
+  },
+  {
+    "id": "100015319",
+    "uri": "http://dx.doi.org/10.13039/100015319",
+    "name": "Sport Ireland Institute",
+    "synonym": []
+  },
+  {
+    "id": "100015320",
+    "uri": "http://dx.doi.org/10.13039/100015320",
+    "name": "Paralympics Ireland",
+    "synonym": []
+  },
+  {
+    "id": "100015442",
+    "uri": "http://dx.doi.org/10.13039/100015442",
+    "name": "Global Brain Health Institute",
+    "synonym": []
+  },
+  {
+    "id": "100015776",
+    "uri": "http://dx.doi.org/10.13039/100015776",
+    "name": "Health and Social Care Board",
+    "synonym": []
+  },
+  {
+    "id": "100015992",
+    "uri": "http://dx.doi.org/10.13039/100015992",
+    "name": "St. Luke's Institute of Cancer Research",
+    "synonym": []
+  },
+  {
+    "id": "100017897",
+    "uri": "http://dx.doi.org/10.13039/100017897",
+    "name": "Friedreich\u2019s Ataxia Research Alliance Ireland",
+    "synonym": []
+  },
+  {
+    "id": "100018064",
+    "uri": "http://dx.doi.org/10.13039/100018064",
+    "name": "Department of Tourism, Culture, Arts, Gaeltacht, Sport and Media",
+    "synonym": []
+  },
+  {
+    "id": "100018172",
+    "uri": "http://dx.doi.org/10.13039/100018172",
+    "name": "Department of the Environment, Climate and Communications",
+    "synonym": []
+  },
+  {
+    "id": "100018175",
+    "uri": "http://dx.doi.org/10.13039/100018175",
+    "name": "Dairy Processing Technology Centre",
+    "synonym": []
+  },
+  {
+    "id": "100018270",
+    "uri": "http://dx.doi.org/10.13039/100018270",
+    "name": "Health Service Executive",
+    "synonym": []
+  },
+  {
+    "id": "100018529",
+    "uri": "http://dx.doi.org/10.13039/100018529",
+    "name": "Alkermes",
+    "synonym": []
+  },
+  {
+    "id": "100018542",
+    "uri": "http://dx.doi.org/10.13039/100018542",
+    "name": "Irish Endocrine Society",
+    "synonym": []
+  },
+  {
+    "id": "100018754",
+    "uri": "http://dx.doi.org/10.13039/100018754",
+    "name": "An Roinn Sl\u00e1inte",
+    "synonym": []
+  },
+  {
+    "id": "100018998",
+    "uri": "http://dx.doi.org/10.13039/100018998",
+    "name": "Irish Research eLibrary",
+    "synonym": []
+  },
+  {
+    "id": "100019428",
+    "uri": "http://dx.doi.org/10.13039/100019428",
+    "name": "Nabriva Therapeutics",
+    "synonym": []
+  },
+  {
+    "id": "100019637",
+    "uri": "http://dx.doi.org/10.13039/100019637",
+    "name": "Horizon Therapeutics",
+    "synonym": []
+  },
+  {
+    "id": "100020174",
+    "uri": "http://dx.doi.org/10.13039/100020174",
+    "name": "Health Research Charities Ireland",
+    "synonym": []
+  },
+  {
+    "id": "100020202",
+    "uri": "http://dx.doi.org/10.13039/100020202",
+    "name": "UCD Foundation",
+    "synonym": []
+  },
+  {
+    "id": "100020233",
+    "uri": "http://dx.doi.org/10.13039/100020233",
+    "name": "Ireland Canada University Foundation",
+    "synonym": []
+  },
+  {
+    "id": "100022943",
+    "uri": "http://dx.doi.org/10.13039/100022943",
+    "name": "National Cancer Registry Ireland",
+    "synonym": []
+  },
+  {
+    "id": "501100001581",
+    "uri": "http://dx.doi.org/10.13039/501100001581",
+    "name": "Arts Council of Ireland",
+    "synonym": []
+  },
+  {
+    "id": "501100001582",
+    "uri": "http://dx.doi.org/10.13039/501100001582",
+    "name": "Centre for Ageing Research and Development in Ireland",
+    "synonym": []
+  },
+  {
+    "id": "501100001583",
+    "uri": "http://dx.doi.org/10.13039/501100001583",
+    "name": "Cystinosis Foundation Ireland",
+    "synonym": []
+  },
+  {
+    "id": "501100001584",
+    "uri": "http://dx.doi.org/10.13039/501100001584",
+    "name": "Department of Agriculture, Food and the Marine, Ireland",
+    "synonym": []
+  },
+  {
+    "id": "501100001586",
+    "uri": "http://dx.doi.org/10.13039/501100001586",
+    "name": "Department of Education and Skills, Ireland",
+    "synonym": []
+  },
+  {
+    "id": "501100001587",
+    "uri": "http://dx.doi.org/10.13039/501100001587",
+    "name": "Economic and Social Research Institute",
+    "synonym": []
+  },
+  {
+    "id": "501100001588",
+    "uri": "http://dx.doi.org/10.13039/501100001588",
+    "name": "Enterprise Ireland",
+    "synonym": []
+  },
+  {
+    "id": "501100001589",
+    "uri": "http://dx.doi.org/10.13039/501100001589",
+    "name": "Environmental Protection Agency",
+    "synonym": []
+  },
+  {
+    "id": "501100001591",
+    "uri": "http://dx.doi.org/10.13039/501100001591",
+    "name": "Heritage Council",
+    "synonym": []
+  },
+  {
+    "id": "501100001592",
+    "uri": "http://dx.doi.org/10.13039/501100001592",
+    "name": "Higher Education Authority",
+    "synonym": []
+  },
+  {
+    "id": "501100001593",
+    "uri": "http://dx.doi.org/10.13039/501100001593",
+    "name": "Irish Cancer Society",
+    "synonym": []
+  },
+  {
+    "id": "501100001594",
+    "uri": "http://dx.doi.org/10.13039/501100001594",
+    "name": "Irish Heart Foundation",
+    "synonym": []
+  },
+  {
+    "id": "501100001595",
+    "uri": "http://dx.doi.org/10.13039/501100001595",
+    "name": "Irish Hospice Foundation",
+    "synonym": []
+  },
+  {
+    "id": "501100001596",
+    "uri": "http://dx.doi.org/10.13039/501100001596",
+    "name": "Irish Research Council for Science, Engineering and Technology",
+    "synonym": []
+  },
+  {
+    "id": "501100001597",
+    "uri": "http://dx.doi.org/10.13039/501100001597",
+    "name": "Irish Research Council for the Humanities and Social Sciences",
+    "synonym": []
+  },
+  {
+    "id": "501100001598",
+    "uri": "http://dx.doi.org/10.13039/501100001598",
+    "name": "Mental Health Commission",
+    "synonym": []
+  },
+  {
+    "id": "501100001600",
+    "uri": "http://dx.doi.org/10.13039/501100001600",
+    "name": "Research and Education Foundation, Sligo General Hospital",
+    "synonym": []
+  },
+  {
+    "id": "501100001601",
+    "uri": "http://dx.doi.org/10.13039/501100001601",
+    "name": "Royal Irish Academy",
+    "synonym": []
+  },
+  {
+    "id": "501100001603",
+    "uri": "http://dx.doi.org/10.13039/501100001603",
+    "name": "Sustainable Energy Authority of Ireland",
+    "synonym": []
+  },
+  {
+    "id": "501100001604",
+    "uri": "http://dx.doi.org/10.13039/501100001604",
+    "name": "Teagasc",
+    "synonym": []
+  },
+  {
+    "id": "501100001627",
+    "uri": "http://dx.doi.org/10.13039/501100001627",
+    "name": "Marine Institute",
+    "synonym": []
+  },
+  {
+    "id": "501100001628",
+    "uri": "http://dx.doi.org/10.13039/501100001628",
+    "name": "Central Remedial Clinic",
+    "synonym": []
+  },
+  {
+    "id": "501100001629",
+    "uri": "http://dx.doi.org/10.13039/501100001629",
+    "name": "Royal Dublin Society",
+    "synonym": []
+  },
+  {
+    "id": "501100001630",
+    "uri": "http://dx.doi.org/10.13039/501100001630",
+    "name": "Dublin Institute for Advanced Studies",
+    "synonym": []
+  },
+  {
+    "id": "501100001631",
+    "uri": "http://dx.doi.org/10.13039/501100001631",
+    "name": "University College Dublin",
+    "synonym": []
+  },
+  {
+    "id": "501100001633",
+    "uri": "http://dx.doi.org/10.13039/501100001633",
+    "name": "National University of Ireland, Maynooth",
+    "synonym": []
+  },
+  {
+    "id": "501100001634",
+    "uri": "http://dx.doi.org/10.13039/501100001634",
+    "name": "University of Galway",
+    "synonym": []
+  },
+  {
+    "id": "501100001635",
+    "uri": "http://dx.doi.org/10.13039/501100001635",
+    "name": "University of Limerick",
+    "synonym": []
+  },
+  {
+    "id": "501100001636",
+    "uri": "http://dx.doi.org/10.13039/501100001636",
+    "name": "University College Cork",
+    "synonym": []
+  },
+  {
+    "id": "501100001637",
+    "uri": "http://dx.doi.org/10.13039/501100001637",
+    "name": "Trinity College Dublin",
+    "synonym": []
+  },
+  {
+    "id": "501100001638",
+    "uri": "http://dx.doi.org/10.13039/501100001638",
+    "name": "Dublin City University",
+    "synonym": []
+  },
+  {
+    "id": "501100002081",
+    "uri": "http://dx.doi.org/10.13039/501100002081",
+    "name": "Irish Research Council",
+    "synonym": []
+  },
+  {
+    "id": "501100002736",
+    "uri": "http://dx.doi.org/10.13039/501100002736",
+    "name": "Covidien",
+    "synonym": []
+  },
+  {
+    "id": "501100002755",
+    "uri": "http://dx.doi.org/10.13039/501100002755",
+    "name": "Brennan and Company",
+    "synonym": []
+  },
+  {
+    "id": "501100002919",
+    "uri": "http://dx.doi.org/10.13039/501100002919",
+    "name": "Cork Institute of Technology",
+    "synonym": []
+  },
+  {
+    "id": "501100002959",
+    "uri": "http://dx.doi.org/10.13039/501100002959",
+    "name": "Dublin City Council",
+    "synonym": []
+  },
+  {
+    "id": "501100003036",
+    "uri": "http://dx.doi.org/10.13039/501100003036",
+    "name": "Perrigo Company Charitable Foundation",
+    "synonym": []
+  },
+  {
+    "id": "501100003037",
+    "uri": "http://dx.doi.org/10.13039/501100003037",
+    "name": "Elan",
+    "synonym": []
+  },
+  {
+    "id": "501100003496",
+    "uri": "http://dx.doi.org/10.13039/501100003496",
+    "name": "HeyStaks Technologies",
+    "synonym": []
+  },
+  {
+    "id": "501100003553",
+    "uri": "http://dx.doi.org/10.13039/501100003553",
+    "name": "Gaelic Athletic Association",
+    "synonym": []
+  },
+  {
+    "id": "501100003840",
+    "uri": "http://dx.doi.org/10.13039/501100003840",
+    "name": "Irish Institute of Clinical Neuroscience",
+    "synonym": []
+  },
+  {
+    "id": "501100003956",
+    "uri": "http://dx.doi.org/10.13039/501100003956",
+    "name": "Aspect Medical Systems",
+    "synonym": []
+  },
+  {
+    "id": "501100004162",
+    "uri": "http://dx.doi.org/10.13039/501100004162",
+    "name": "Meath Foundation",
+    "synonym": []
+  },
+  {
+    "id": "501100004210",
+    "uri": "http://dx.doi.org/10.13039/501100004210",
+    "name": "Our Lady's Children's Hospital, Crumlin",
+    "synonym": []
+  },
+  {
+    "id": "501100004321",
+    "uri": "http://dx.doi.org/10.13039/501100004321",
+    "name": "Shire",
+    "synonym": []
+  },
+  {
+    "id": "501100004981",
+    "uri": "http://dx.doi.org/10.13039/501100004981",
+    "name": "Athlone Institute of Technology",
+    "synonym": []
+  },
+  {
+    "id": "501100006518",
+    "uri": "http://dx.doi.org/10.13039/501100006518",
+    "name": "Department of Communications, Energy and Natural Resources, Ireland",
+    "synonym": []
+  },
+  {
+    "id": "501100006553",
+    "uri": "http://dx.doi.org/10.13039/501100006553",
+    "name": "Collaborative Centre for Applied Nanotechnology",
+    "synonym": []
+  },
+  {
+    "id": "501100006759",
+    "uri": "http://dx.doi.org/10.13039/501100006759",
+    "name": "CLARITY Centre for Sensor Web Technologies",
+    "synonym": []
+  },
+  {
+    "id": "501100009246",
+    "uri": "http://dx.doi.org/10.13039/501100009246",
+    "name": "Technological University Dublin",
+    "synonym": []
+  },
+  {
+    "id": "501100009269",
+    "uri": "http://dx.doi.org/10.13039/501100009269",
+    "name": "Programme of Competitive Forestry Research for Development",
+    "synonym": []
+  },
+  {
+    "id": "501100009315",
+    "uri": "http://dx.doi.org/10.13039/501100009315",
+    "name": "Cystinosis Ireland",
+    "synonym": []
+  },
+  {
+    "id": "501100010808",
+    "uri": "http://dx.doi.org/10.13039/501100010808",
+    "name": "Geological Survey of Ireland",
+    "synonym": []
+  },
+  {
+    "id": "501100011030",
+    "uri": "http://dx.doi.org/10.13039/501100011030",
+    "name": "Alimentary Glycoscience Research Cluster",
+    "synonym": []
+  },
+  {
+    "id": "501100011031",
+    "uri": "http://dx.doi.org/10.13039/501100011031",
+    "name": "Alimentary Health",
+    "synonym": []
+  },
+  {
+    "id": "501100011103",
+    "uri": "http://dx.doi.org/10.13039/501100011103",
+    "name": "Rann\u00eds",
+    "synonym": []
+  },
+  {
+    "id": "501100012354",
+    "uri": "http://dx.doi.org/10.13039/501100012354",
+    "name": "Inland Fisheries Ireland",
+    "synonym": []
+  },
+  {
+    "id": "501100014384",
+    "uri": "http://dx.doi.org/10.13039/501100014384",
+    "name": "X-Bolt Orthopaedics",
+    "synonym": []
+  },
+  {
+    "id": "501100014710",
+    "uri": "http://dx.doi.org/10.13039/501100014710",
+    "name": "PrecisionBiotics Group",
+    "synonym": []
+  },
+  {
+    "id": "501100014827",
+    "uri": "http://dx.doi.org/10.13039/501100014827",
+    "name": "Dormant Accounts Fund",
+    "synonym": []
+  },
+  {
+    "id": "501100016041",
+    "uri": "http://dx.doi.org/10.13039/501100016041",
+    "name": "St Vincents Anaesthesia Foundation",
+    "synonym": []
+  },
+  {
+    "id": "501100017501",
+    "uri": "http://dx.doi.org/10.13039/501100017501",
+    "name": "FotoNation",
+    "synonym": []
+  },
+  {
+    "id": "501100018641",
+    "uri": "http://dx.doi.org/10.13039/501100018641",
+    "name": "Dairy Research Ireland",
+    "synonym": []
+  },
+  {
+    "id": "501100018839",
+    "uri": "http://dx.doi.org/10.13039/501100018839",
+    "name": "Irish Centre for High-End Computing",
+    "synonym": []
+  },
+  {
+    "id": "501100019905",
+    "uri": "http://dx.doi.org/10.13039/501100019905",
+    "name": "Galway University Foundation",
+    "synonym": []
+  },
+  {
+    "id": "501100020036",
+    "uri": "http://dx.doi.org/10.13039/501100020036",
+    "name": "Dystonia Ireland",
+    "synonym": []
+  },
+  {
+    "id": "501100020221",
+    "uri": "http://dx.doi.org/10.13039/501100020221",
+    "name": "Irish Motor Neurone Disease Association",
+    "synonym": []
+  },
+  {
+    "id": "501100020270",
+    "uri": "http://dx.doi.org/10.13039/501100020270",
+    "name": "Advanced Materials and Bioengineering Research",
+    "synonym": []
+  },
+  {
+    "id": "501100020403",
+    "uri": "http://dx.doi.org/10.13039/501100020403",
+    "name": "Irish Composites Centre",
+    "synonym": []
+  },
+  {
+    "id": "501100020425",
+    "uri": "http://dx.doi.org/10.13039/501100020425",
+    "name": "Irish Thoracic Society",
+    "synonym": []
+  },
+  {
+    "id": "501100021102",
+    "uri": "http://dx.doi.org/10.13039/501100021102",
+    "name": "Waterford Institute of Technology",
+    "synonym": []
+  },
+  {
+    "id": "501100021110",
+    "uri": "http://dx.doi.org/10.13039/501100021110",
+    "name": "Irish MPS Society",
+    "synonym": []
+  },
+  {
+    "id": "501100021525",
+    "uri": "http://dx.doi.org/10.13039/501100021525",
+    "name": "Insight SFI Research Centre for Data Analytics",
+    "synonym": []
+  },
+  {
+    "id": "501100021694",
+    "uri": "http://dx.doi.org/10.13039/501100021694",
+    "name": "Elan Pharma International",
+    "synonym": []
+  },
+  {
+    "id": "501100021838",
+    "uri": "http://dx.doi.org/10.13039/501100021838",
+    "name": "Royal College of Physicians of Ireland",
+    "synonym": []
+  },
+  {
+    "id": "501100022542",
+    "uri": "http://dx.doi.org/10.13039/501100022542",
+    "name": "Breakthrough Cancer Research",
+    "synonym": []
+  },
+  {
+    "id": "501100022610",
+    "uri": "http://dx.doi.org/10.13039/501100022610",
+    "name": "Breast Cancer Ireland",
+    "synonym": []
+  },
+  {
+    "id": "501100022728",
+    "uri": "http://dx.doi.org/10.13039/501100022728",
+    "name": "Munster Technological University",
+    "synonym": []
+  },
+  {
+    "id": "501100022729",
+    "uri": "http://dx.doi.org/10.13039/501100022729",
+    "name": "Institute of Technology, Tralee",
+    "synonym": []
+  },
+  {
+    "id": "501100023273",
+    "uri": "http://dx.doi.org/10.13039/501100023273",
+    "name": "HRB Clinical Research Facility Galway",
+    "synonym": []
+  },
+  {
+    "id": "501100023378",
+    "uri": "http://dx.doi.org/10.13039/501100023378",
+    "name": "Lauritzson Foundation",
+    "synonym": []
+  },
+  {
+    "id": "501100023551",
+    "uri": "http://dx.doi.org/10.13039/501100023551",
+    "name": "Cystic Fibrosis Ireland",
+    "synonym": []
+  },
+  {
+    "id": "501100023970",
+    "uri": "http://dx.doi.org/10.13039/501100023970",
+    "name": "Tyndall National Institute",
+    "synonym": []
+  },
+  {
+    "id": "501100024094",
+    "uri": "http://dx.doi.org/10.13039/501100024094",
+    "name": "Raidi\u00f3 Teilif\u00eds \u00c9ireann",
+    "synonym": []
+  },
+  {
+    "id": "501100024242",
+    "uri": "http://dx.doi.org/10.13039/501100024242",
+    "name": "Synthesis and Solid State Pharmaceutical Centre",
+    "synonym": []
+  },
+  {
+    "id": "501100024313",
+    "uri": "http://dx.doi.org/10.13039/501100024313",
+    "name": "Irish Rugby Football Union",
+    "synonym": []
+  },
+  {
+    "id": "100007490",
+    "uri": "http://dx.doi.org/10.13039/100007490",
+    "name": "Bausch and Lomb Ireland",
+    "synonym": []
+  },
+  {
+    "id": "100007819",
+    "uri": "http://dx.doi.org/10.13039/100007819",
+    "name": "Allergan",
+    "synonym": []
+  },
+  {
+    "id": "100010547",
+    "uri": "http://dx.doi.org/10.13039/100010547",
+    "name": "Irish Youth Justice Service",
+    "synonym": []
+  },
+  {
+    "id": "100012733",
+    "uri": "http://dx.doi.org/10.13039/100012733",
+    "name": "National Parks and Wildlife Service",
+    "synonym": []
+  },
+  {
+    "id": "100015278",
+    "uri": "http://dx.doi.org/10.13039/100015278",
+    "name": "Pfizer Healthcare Ireland",
+    "synonym": []
+  },
+  {
+    "id": "100017144",
+    "uri": "http://dx.doi.org/10.13039/100017144",
+    "name": "Shell E and P Ireland",
+    "synonym": []
+  },
+  {
+    "id": "100022895",
+    "uri": "http://dx.doi.org/10.13039/100022895",
+    "name": "Health Research Institute, University of Limerick",
+    "synonym": []
+  },
+  {
+    "id": "501100001599",
+    "uri": "http://dx.doi.org/10.13039/501100001599",
+    "name": "National Council for Forest Research and Development",
+    "synonym": []
+  },
+  {
+    "id": "501100006554",
+    "uri": "http://dx.doi.org/10.13039/501100006554",
+    "name": "IDA Ireland",
+    "synonym": []
+  },
+  {
+    "id": "501100011626",
+    "uri": "http://dx.doi.org/10.13039/501100011626",
+    "name": "Energy Policy Research Centre, Economic and Social Research Institute",
+    "synonym": []
+  },
+  {
+    "id": "501100014531",
+    "uri": "http://dx.doi.org/10.13039/501100014531",
+    "name": "Physical Education and Sport Sciences Department, University of Limerick",
+    "synonym": []
+  },
+  {
+    "id": "501100014745",
+    "uri": "http://dx.doi.org/10.13039/501100014745",
+    "name": "APC Microbiome Institute",
+    "synonym": []
+  },
+  {
+    "id": "501100014826",
+    "uri": "http://dx.doi.org/10.13039/501100014826",
+    "name": "ADAPT - Centre for Digital Content Technology",
+    "synonym": []
+  },
+  {
+    "id": "501100020570",
+    "uri": "http://dx.doi.org/10.13039/501100020570",
+    "name": "College of Medicine, Nursing and Health Sciences, National University of Ireland, Galway",
+    "synonym": []
+  },
+  {
+    "id": "501100020871",
+    "uri": "http://dx.doi.org/10.13039/501100020871",
+    "name": "Bernal Institute, University of Limerick",
+    "synonym": []
+  },
+  {
+    "id": "501100023852",
+    "uri": "http://dx.doi.org/10.13039/501100023852",
+    "name": "Moore Institute for Research in the Humanities and Social Studies, University of Galway",
+    "synonym": []
+  }
+]
--- a/dhp-workflows/dhp-doiboost/src/main/scala/eu/dnetlib/doiboost/crossref/Crossref2Oaf.scala
+++ b/dhp-workflows/dhp-doiboost/src/main/scala/eu/dnetlib/doiboost/crossref/Crossref2Oaf.scala
@ -16,6 +16,7 @@ import org.slf4j.{Logger, LoggerFactory}
 import java.util
 import scala.collection.JavaConverters._
 import scala.collection.mutable
+import scala.io.Source
 import scala.util.matching.Regex

 case class CrossrefDT(doi: String, json: String, timestamp: Long) {}
@ -30,11 +31,22 @@ case class mappingAuthor(
  affiliation: Option[mappingAffiliation]
 ) {}

+case class funderInfo(id: String, uri: String, name: String, synonym: List[String]) {}
+
 case class mappingFunder(name: String, DOI: Option[String], award: Option[List[String]]) {}

 case object Crossref2Oaf {
  val logger: Logger = LoggerFactory.getLogger(Crossref2Oaf.getClass)

+  val irishFunder: List[funderInfo] = {
+    val s = Source
+      .fromInputStream(getClass.getResourceAsStream("/eu/dnetlib/dhp/doiboost/crossref/irish_funder.json"))
+      .mkString
+    implicit lazy val formats: DefaultFormats.type = org.json4s.DefaultFormats
+    lazy val json: org.json4s.JValue = parse(s)
+    json.extract[List[funderInfo]]
+  }
+
  val mappingCrossrefType = Map(
    "book-section"        -> "publication",
    "book"                -> "publication",
@ -88,6 +100,13 @@ case object Crossref2Oaf {
    "report"              -> "0017 Report"
  )

+  def getIrishId(doi: String): Option[String] = {
+    val id = doi.split("/").last
+    irishFunder
+      .find(f => id.equalsIgnoreCase(f.id) || (f.synonym.nonEmpty && f.synonym.exists(s => s.equalsIgnoreCase(id))))
+      .map(f => f.id)
+  }
+
  def mappingResult(result: Result, json: JValue, cobjCategory: String): Result = {
    implicit lazy val formats: DefaultFormats.type = org.json4s.DefaultFormats

@ -467,6 +486,14 @@ case object Crossref2Oaf {
    if (funders != null)
      funders.foreach(funder => {
        if (funder.DOI.isDefined && funder.DOI.get.nonEmpty) {
+
+          if (getIrishId(funder.DOI.get).isDefined) {
+            val nsPrefix = getIrishId(funder.DOI.get).get.padTo(12, '_')
+            val targetId = getProjectId(nsPrefix, "1e5e62235d094afd01cd56e65112fc63")
+            queue += generateRelation(sourceId, targetId, ModelConstants.IS_PRODUCED_BY)
+            queue += generateRelation(targetId, sourceId, ModelConstants.PRODUCES)
+          }
+
          funder.DOI.get match {
            case "10.13039/100010663" | "10.13039/100010661" | "10.13039/501100007601" | "10.13039/501100000780" |
                "10.13039/100010665" =>
--- a/dhp-workflows/dhp-enrichment/src/main/java/eu/dnetlib/dhp/PropagationConstant.java
+++ b/dhp-workflows/dhp-enrichment/src/main/java/eu/dnetlib/dhp/PropagationConstant.java
@ -71,6 +71,9 @@ public class PropagationConstant {
 	public static final String PROPAGATION_RESULT_COMMUNITY_ORGANIZATION_CLASS_ID = "result:community:organization";
 	public static final String PROPAGATION_RESULT_COMMUNITY_ORGANIZATION_CLASS_NAME = " Propagation of result belonging to community through organization";

+	public static final String PROPAGATION_RESULT_COMMUNITY_PROJECT_CLASS_ID = "result:community:project";
+	public static final String PROPAGATION_RESULT_COMMUNITY_PROJECT_CLASS_NAME = " Propagation of result belonging to community through project";
+
 	public static final String PROPAGATION_ORCID_TO_RESULT_FROM_SEM_REL_CLASS_ID = "authorpid:result";
 	public static final String PROPAGATION_ORCID_TO_RESULT_FROM_SEM_REL_CLASS_NAME = "Propagation of authors pid to result through semantic relations";

--- a/dhp-workflows/dhp-enrichment/src/main/java/eu/dnetlib/dhp/api/QueryCommunityAPI.java
+++ b/dhp-workflows/dhp-enrichment/src/main/java/eu/dnetlib/dhp/api/QueryCommunityAPI.java
@ -0,0 +1,83 @@
+
+package eu.dnetlib.dhp.api;
+
+import java.io.BufferedReader;
+import java.io.IOException;
+import java.io.InputStreamReader;
+import java.net.HttpURLConnection;
+import java.net.URL;
+
+import org.jetbrains.annotations.NotNull;
+
+/**
+ * @author miriam.baglioni
+ * @Date 06/10/23
+ */
+public class QueryCommunityAPI {
+	private static final String PRODUCTION_BASE_URL = "https://services.openaire.eu/openaire/";
+	private static final String BETA_BASE_URL = "https://beta.services.openaire.eu/openaire/";
+
+	private static String get(String geturl) throws IOException {
+		URL url = new URL(geturl);
+		HttpURLConnection conn = (HttpURLConnection) url.openConnection();
+		conn.setDoOutput(true);
+		conn.setRequestMethod("GET");
+
+		int responseCode = conn.getResponseCode();
+		String body = getBody(conn);
+		conn.disconnect();
+		if (responseCode != HttpURLConnection.HTTP_OK)
+			throw new IOException("Unexpected code " + responseCode + body);
+
+		return body;
+	}
+
+	public static String communities(boolean production) throws IOException {
+		if (production)
+			return get(PRODUCTION_BASE_URL + "community/communities");
+		return get(BETA_BASE_URL + "community/communities");
+	}
+
+	public static String community(String id, boolean production) throws IOException {
+		if (production)
+			return get(PRODUCTION_BASE_URL + "community/" + id);
+		return get(BETA_BASE_URL + "community/" + id);
+	}
+
+	public static String communityDatasource(String id, boolean production) throws IOException {
+		if (production)
+			return get(PRODUCTION_BASE_URL + "community/" + id + "/contentproviders");
+		return (BETA_BASE_URL + "community/" + id + "/contentproviders");
+
+	}
+
+	public static String communityPropagationOrganization(String id, boolean production) throws IOException {
+		if (production)
+			return get(PRODUCTION_BASE_URL + "community/" + id + "/propagationOrganizations");
+		return get(BETA_BASE_URL + "community/" + id + "/propagationOrganizations");
+	}
+
+	public static String communityProjects(String id, String page, String size, boolean production) throws IOException {
+		if (production)
+			return get(PRODUCTION_BASE_URL + "community/" + id + "/projects/" + page + "/" + size);
+		return get(BETA_BASE_URL + "community/" + id + "/projects/" + page + "/" + size);
+	}
+
+	@NotNull
+	private static String getBody(HttpURLConnection conn) throws IOException {
+		String body = "{}";
+		try (BufferedReader br = new BufferedReader(
+			new InputStreamReader(conn.getInputStream(), "utf-8"))) {
+			StringBuilder response = new StringBuilder();
+			String responseLine = null;
+			while ((responseLine = br.readLine()) != null) {
+				response.append(responseLine.trim());
+			}
+
+			body = response.toString();
+
+		}
+		return body;
+	}
+
+}
--- a/dhp-workflows/dhp-enrichment/src/main/java/eu/dnetlib/dhp/api/Utils.java
+++ b/dhp-workflows/dhp-enrichment/src/main/java/eu/dnetlib/dhp/api/Utils.java
@ -0,0 +1,169 @@
+
+package eu.dnetlib.dhp.api;
+
+import java.io.IOException;
+import java.io.Serializable;
+import java.util.ArrayList;
+import java.util.List;
+import java.util.Map;
+import java.util.Objects;
+import java.util.stream.Collectors;
+
+import javax.management.Query;
+
+import org.slf4j.Logger;
+import org.slf4j.LoggerFactory;
+
+import com.amazonaws.util.StringUtils;
+import com.fasterxml.jackson.databind.ObjectMapper;
+import com.google.common.collect.Maps;
+
+import eu.dnetlib.dhp.api.model.*;
+import eu.dnetlib.dhp.bulktag.community.Community;
+import eu.dnetlib.dhp.bulktag.community.CommunityConfiguration;
+import eu.dnetlib.dhp.bulktag.community.Provider;
+import eu.dnetlib.dhp.bulktag.criteria.VerbResolver;
+import eu.dnetlib.dhp.bulktag.criteria.VerbResolverFactory;
+import eu.dnetlib.dhp.resulttocommunityfromorganization.SparkResultToCommunityFromOrganizationJob;
+
+/**
+ * @author miriam.baglioni
+ * @Date 09/10/23
+ */
+public class Utils implements Serializable {
+	private static final ObjectMapper MAPPER = new ObjectMapper();
+	private static final VerbResolver resolver = VerbResolverFactory.newInstance();
+
+	private static final Logger log = LoggerFactory.getLogger(Utils.class);
+
+	public static CommunityConfiguration getCommunityConfiguration(boolean production) throws IOException {
+		final Map<String, Community> communities = Maps.newHashMap();
+		List<Community> validCommunities = new ArrayList<>();
+		getValidCommunities(production)
+			.forEach(community -> {
+				try {
+					CommunityModel cm = MAPPER
+						.readValue(QueryCommunityAPI.community(community.getId(), production), CommunityModel.class);
+					validCommunities.add(getCommunity(cm));
+				} catch (IOException e) {
+					throw new RuntimeException(e);
+				}
+			});
+		validCommunities.forEach(community -> {
+			try {
+				DatasourceList dl = MAPPER
+					.readValue(
+						QueryCommunityAPI.communityDatasource(community.getId(), production), DatasourceList.class);
+				community.setProviders(dl.stream().map(d -> {
+					if (d.getEnabled() == null || Boolean.FALSE.equals(d.getEnabled()))
+						return null;
+					Provider p = new Provider();
+					p.setOpenaireId("10|" + d.getOpenaireId());
+					p.setSelectionConstraints(d.getSelectioncriteria());
+					if (p.getSelectionConstraints() != null)
+						p.getSelectionConstraints().setSelection(resolver);
+					return p;
+				})
+					.filter(Objects::nonNull)
+					.collect(Collectors.toList()));
+			} catch (IOException e) {
+				throw new RuntimeException(e);
+			}
+		});
+
+		validCommunities.forEach(community -> {
+			if (community.isValid())
+				communities.put(community.getId(), community);
+		});
+		return new CommunityConfiguration(communities);
+	}
+
+	private static Community getCommunity(CommunityModel cm) {
+		Community c = new Community();
+		c.setId(cm.getId());
+		c.setZenodoCommunities(cm.getOtherZenodoCommunities());
+		if (!StringUtils.isNullOrEmpty(cm.getZenodoCommunity()))
+			c.getZenodoCommunities().add(cm.getZenodoCommunity());
+		c.setSubjects(cm.getSubjects());
+		c.getSubjects().addAll(cm.getFos());
+		c.getSubjects().addAll(cm.getSdg());
+		if (cm.getAdvancedConstraints() != null) {
+			c.setConstraints(cm.getAdvancedConstraints());
+			c.getConstraints().setSelection(resolver);
+		}
+		if (cm.getRemoveConstraints() != null) {
+			c.setRemoveConstraints(cm.getRemoveConstraints());
+			c.getRemoveConstraints().setSelection(resolver);
+		}
+		return c;
+	}
+
+	public static List<CommunityModel> getValidCommunities(boolean production) throws IOException {
+		return MAPPER
+			.readValue(QueryCommunityAPI.communities(production), CommunitySummary.class)
+			.stream()
+			.filter(
+				community -> !community.getStatus().equals("hidden") &&
+					(community.getType().equals("ri") || community.getType().equals("community")))
+			.collect(Collectors.toList());
+	}
+
+	/**
+	 * it returns for each organization the list of associated communities
+	 */
+	public static CommunityEntityMap getCommunityOrganization(boolean production) throws IOException {
+		CommunityEntityMap organizationMap = new CommunityEntityMap();
+		getValidCommunities(production)
+			.forEach(community -> {
+				String id = community.getId();
+				try {
+					List<String> associatedOrgs = MAPPER
+						.readValue(
+							QueryCommunityAPI.communityPropagationOrganization(id, production), OrganizationList.class);
+					associatedOrgs.forEach(o -> {
+						if (!organizationMap
+							.keySet()
+							.contains(
+								"20|" + o))
+							organizationMap.put("20|" + o, new ArrayList<>());
+						organizationMap.get("20|" + o).add(community.getId());
+					});
+				} catch (IOException e) {
+					throw new RuntimeException(e);
+				}
+			});
+
+		return organizationMap;
+	}
+
+	public static CommunityEntityMap getCommunityProjects(boolean production) throws IOException {
+		CommunityEntityMap projectMap = new CommunityEntityMap();
+		getValidCommunities(production)
+			.forEach(community -> {
+				int page = -1;
+				int size = 100;
+				ContentModel cm = new ContentModel();
+				do {
+					page++;
+					try {
+						cm = MAPPER
+							.readValue(
+								QueryCommunityAPI
+									.communityProjects(
+										community.getId(), String.valueOf(page), String.valueOf(size), production),
+								ContentModel.class);
+						if (cm.getContent().size() > 0) {
+							cm.getContent().forEach(p -> {
+								if (!projectMap.keySet().contains("40|" + p.getOpenaireId()))
+									projectMap.put("40|" + p.getOpenaireId(), new ArrayList<>());
+								projectMap.get("40|" + p.getOpenaireId()).add(community.getId());
+							});
+						}
+					} catch (IOException e) {
+						throw new RuntimeException(e);
+					}
+				} while (!cm.getLast());
+			});
+		return projectMap;
+	}
+}
--- a/dhp-workflows/dhp-enrichment/src/main/java/eu/dnetlib/dhp/api/model/CommunityContentprovider.java
+++ b/dhp-workflows/dhp-enrichment/src/main/java/eu/dnetlib/dhp/api/model/CommunityContentprovider.java
@ -0,0 +1,43 @@
+
+package eu.dnetlib.dhp.api.model;
+
+import com.fasterxml.jackson.annotation.JsonAutoDetect;
+import com.fasterxml.jackson.annotation.JsonIgnoreProperties;
+import com.google.gson.Gson;
+
+import eu.dnetlib.dhp.bulktag.community.SelectionConstraints;
+
+@JsonAutoDetect
+@JsonIgnoreProperties(ignoreUnknown = true)
+public class CommunityContentprovider {
+	private String openaireId;
+	private SelectionConstraints selectioncriteria;
+
+	private String enabled;
+
+	public String getEnabled() {
+		return enabled;
+	}
+
+	public void setEnabled(String enabled) {
+		this.enabled = enabled;
+	}
+
+	public String getOpenaireId() {
+		return openaireId;
+	}
+
+	public void setOpenaireId(final String openaireId) {
+		this.openaireId = openaireId;
+	}
+
+	public SelectionConstraints getSelectioncriteria() {
+
+		return this.selectioncriteria;
+	}
+
+	public void setSelectioncriteria(SelectionConstraints selectioncriteria) {
+		this.selectioncriteria = selectioncriteria;
+
+	}
+}
--- a/dhp-workflows/dhp-enrichment/src/main/java/eu/dnetlib/dhp/resulttocommunityfromorganization/OrganizationMap.java
+++ b/dhp-workflows/dhp-enrichment/src/main/java/eu/dnetlib/dhp/resulttocommunityfromorganization/OrganizationMap.java
@ -1,13 +1,13 @@

-package eu.dnetlib.dhp.resulttocommunityfromorganization;
+package eu.dnetlib.dhp.api.model;

 import java.util.ArrayList;
 import java.util.HashMap;
 import java.util.List;

-public class OrganizationMap extends HashMap<String, List<String>> {
+public class CommunityEntityMap extends HashMap<String, List<String>> {

-	public OrganizationMap() {
+	public CommunityEntityMap() {
 		super();
 	}

--- a/dhp-workflows/dhp-enrichment/src/main/java/eu/dnetlib/dhp/api/model/CommunityModel.java
+++ b/dhp-workflows/dhp-enrichment/src/main/java/eu/dnetlib/dhp/api/model/CommunityModel.java
@ -0,0 +1,108 @@
+
+package eu.dnetlib.dhp.api.model;
+
+import java.io.Serializable;
+import java.util.List;
+
+import com.fasterxml.jackson.annotation.JsonIgnoreProperties;
+
+import eu.dnetlib.dhp.bulktag.community.SelectionConstraints;
+
+/**
+ * @author miriam.baglioni
+ * @Date 06/10/23
+ */
+@JsonIgnoreProperties(ignoreUnknown = true)
+public class CommunityModel implements Serializable {
+	private String id;
+	private String type;
+	private String status;
+
+	private String zenodoCommunity;
+	private List<String> subjects;
+	private List<String> otherZenodoCommunities;
+	private List<String> fos;
+	private List<String> sdg;
+	private SelectionConstraints advancedConstraints;
+	private SelectionConstraints removeConstraints;
+
+	public String getZenodoCommunity() {
+		return zenodoCommunity;
+	}
+
+	public void setZenodoCommunity(String zenodoCommunity) {
+		this.zenodoCommunity = zenodoCommunity;
+	}
+
+	public List<String> getSubjects() {
+		return subjects;
+	}
+
+	public void setSubjects(List<String> subjects) {
+		this.subjects = subjects;
+	}
+
+	public List<String> getOtherZenodoCommunities() {
+		return otherZenodoCommunities;
+	}
+
+	public void setOtherZenodoCommunities(List<String> otherZenodoCommunities) {
+		this.otherZenodoCommunities = otherZenodoCommunities;
+	}
+
+	public List<String> getFos() {
+		return fos;
+	}
+
+	public void setFos(List<String> fos) {
+		this.fos = fos;
+	}
+
+	public List<String> getSdg() {
+		return sdg;
+	}
+
+	public void setSdg(List<String> sdg) {
+		this.sdg = sdg;
+	}
+
+	public SelectionConstraints getRemoveConstraints() {
+		return removeConstraints;
+	}
+
+	public void setRemoveConstraints(SelectionConstraints removeConstraints) {
+		this.removeConstraints = removeConstraints;
+	}
+
+	public SelectionConstraints getAdvancedConstraints() {
+		return advancedConstraints;
+	}
+
+	public void setAdvancedConstraints(SelectionConstraints advancedConstraints) {
+		this.advancedConstraints = advancedConstraints;
+	}
+
+	public String getId() {
+		return id;
+	}
+
+	public void setId(String id) {
+		this.id = id;
+	}
+
+	public String getType() {
+		return type;
+	}
+
+	public void setType(String type) {
+		this.type = type;
+	}
+
+	public String getStatus() {
+		return status;
+	}
+
+	public void setStatus(String status) {
+		this.status = status;
+	}
+}
--- a/dhp-workflows/dhp-enrichment/src/main/java/eu/dnetlib/dhp/api/model/CommunitySummary.java
+++ b/dhp-workflows/dhp-enrichment/src/main/java/eu/dnetlib/dhp/api/model/CommunitySummary.java
@ -0,0 +1,15 @@
+
+package eu.dnetlib.dhp.api.model;
+
+import java.io.Serializable;
+import java.util.ArrayList;
+
+/**
+ * @author miriam.baglioni
+ * @Date 06/10/23
+ */
+public class CommunitySummary extends ArrayList<CommunityModel> implements Serializable {
+	public CommunitySummary() {
+		super();
+	}
+}
--- a/dhp-workflows/dhp-enrichment/src/main/java/eu/dnetlib/dhp/api/model/ContentModel.java
+++ b/dhp-workflows/dhp-enrichment/src/main/java/eu/dnetlib/dhp/api/model/ContentModel.java
@ -0,0 +1,51 @@
+
+package eu.dnetlib.dhp.api.model;
+
+import java.io.Serializable;
+import java.util.List;
+
+import com.fasterxml.jackson.annotation.JsonIgnoreProperties;
+
+/**
+ * @author miriam.baglioni
+ * @Date 09/10/23
+ */
+@JsonIgnoreProperties(ignoreUnknown = true)
+public class ContentModel implements Serializable {
+	private List<ProjectModel> content;
+	private Integer totalPages;
+	private Boolean last;
+	private Integer number;
+
+	public List<ProjectModel> getContent() {
+		return content;
+	}
+
+	public void setContent(List<ProjectModel> content) {
+		this.content = content;
+	}
+
+	public Integer getTotalPages() {
+		return totalPages;
+	}
+
+	public void setTotalPages(Integer totalPages) {
+		this.totalPages = totalPages;
+	}
+
+	public Boolean getLast() {
+		return last;
+	}
+
+	public void setLast(Boolean last) {
+		this.last = last;
+	}
+
+	public Integer getNumber() {
+		return number;
+	}
+
+	public void setNumber(Integer number) {
+		this.number = number;
+	}
+}
--- a/dhp-workflows/dhp-enrichment/src/main/java/eu/dnetlib/dhp/api/model/DatasourceList.java
+++ b/dhp-workflows/dhp-enrichment/src/main/java/eu/dnetlib/dhp/api/model/DatasourceList.java
@ -0,0 +1,13 @@
+
+package eu.dnetlib.dhp.api.model;
+
+import java.io.Serializable;
+import java.util.ArrayList;
+
+import eu.dnetlib.dhp.api.model.CommunityContentprovider;
+
+public class DatasourceList extends ArrayList<CommunityContentprovider> implements Serializable {
+	public DatasourceList() {
+		super();
+	}
+}
--- a/dhp-workflows/dhp-enrichment/src/main/java/eu/dnetlib/dhp/api/model/OrganizationList.java
+++ b/dhp-workflows/dhp-enrichment/src/main/java/eu/dnetlib/dhp/api/model/OrganizationList.java
@ -0,0 +1,16 @@
+
+package eu.dnetlib.dhp.api.model;
+
+import java.io.Serializable;
+import java.util.ArrayList;
+
+/**
+ * @author miriam.baglioni
+ * @Date 09/10/23
+ */
+public class OrganizationList extends ArrayList<String> implements Serializable {
+
+	public OrganizationList() {
+		super();
+	}
+}
--- a/dhp-workflows/dhp-enrichment/src/main/java/eu/dnetlib/dhp/api/model/ProjectModel.java
+++ b/dhp-workflows/dhp-enrichment/src/main/java/eu/dnetlib/dhp/api/model/ProjectModel.java
@ -0,0 +1,24 @@
+
+package eu.dnetlib.dhp.api.model;
+
+import java.io.Serializable;
+
+import com.fasterxml.jackson.annotation.JsonIgnoreProperties;
+
+/**
+ * @author miriam.baglioni
+ * @Date 09/10/23
+ */
+@JsonIgnoreProperties(ignoreUnknown = true)
+public class ProjectModel implements Serializable {
+
+	private String openaireId;
+
+	public String getOpenaireId() {
+		return openaireId;
+	}
+
+	public void setOpenaireId(String openaireId) {
+		this.openaireId = openaireId;
+	}
+}
--- a/dhp-workflows/dhp-enrichment/src/main/java/eu/dnetlib/dhp/bulktag/SparkBulkTagJob.java
+++ b/dhp-workflows/dhp-enrichment/src/main/java/eu/dnetlib/dhp/bulktag/SparkBulkTagJob.java
@ -9,7 +9,6 @@ import java.util.*;
 import org.apache.commons.io.IOUtils;
 import org.apache.spark.SparkConf;
 import org.apache.spark.api.java.function.FilterFunction;
-import org.apache.spark.api.java.function.ForeachFunction;
 import org.apache.spark.api.java.function.MapFunction;
 import org.apache.spark.sql.Dataset;
 import org.apache.spark.sql.Encoders;
@ -21,8 +20,11 @@ import org.slf4j.LoggerFactory;
 import com.fasterxml.jackson.databind.ObjectMapper;
 import com.google.gson.Gson;

+import eu.dnetlib.dhp.api.Utils;
 import eu.dnetlib.dhp.application.ArgumentApplicationParser;
 import eu.dnetlib.dhp.bulktag.community.*;
+import eu.dnetlib.dhp.schema.common.EntityType;
+import eu.dnetlib.dhp.schema.common.ModelSupport;
 import eu.dnetlib.dhp.schema.oaf.Datasource;
 import eu.dnetlib.dhp.schema.oaf.Result;

@ -53,50 +55,38 @@ public class SparkBulkTagJob {
 			.orElse(Boolean.TRUE);
 		log.info("isSparkSessionManaged: {}", isSparkSessionManaged);

-		Boolean isTest = Optional
-			.ofNullable(parser.get("isTest"))
-			.map(Boolean::valueOf)
-			.orElse(Boolean.FALSE);
-		log.info("isTest: {} ", isTest);
-
 		final String inputPath = parser.get("sourcePath");
 		log.info("inputPath: {}", inputPath);

 		final String outputPath = parser.get("outputPath");
 		log.info("outputPath: {}", outputPath);

+		final boolean production = Boolean.valueOf(parser.get("production"));
+		log.info("production: {}", production);
+
 		ProtoMap protoMappingParams = new Gson().fromJson(parser.get("pathMap"), ProtoMap.class);
 		log.info("pathMap: {}", new Gson().toJson(protoMappingParams));

-		final String resultClassName = parser.get("resultTableName");
-		log.info("resultTableName: {}", resultClassName);
-
-		final Boolean saveGraph = Optional
-			.ofNullable(parser.get("saveGraph"))
-			.map(Boolean::valueOf)
-			.orElse(Boolean.TRUE);
-		log.info("saveGraph: {}", saveGraph);
-
-		Class<? extends Result> resultClazz = (Class<? extends Result>) Class.forName(resultClassName);
-
 		SparkConf conf = new SparkConf();
 		CommunityConfiguration cc;

-		String taggingConf = parser.get("taggingConf");
+		String taggingConf = Optional
+			.ofNullable(parser.get("taggingConf"))
+			.map(String::valueOf)
+			.orElse(null);

-		if (isTest) {
+		if (taggingConf != null) {
 			cc = CommunityConfigurationFactory.newInstance(taggingConf);
 		} else {
-			cc = QueryInformationSystem.getCommunityConfiguration(parser.get("isLookUpUrl"));
+			cc = Utils.getCommunityConfiguration(production);
 		}

 		runWithSparkSession(
 			conf,
 			isSparkSessionManaged,
 			spark -> {
-				removeOutputDir(spark, outputPath);
 				extendCommunityConfigurationForEOSC(spark, inputPath, cc);
-				execBulkTag(spark, inputPath, outputPath, protoMappingParams, resultClazz, cc);
+				execBulkTag(spark, inputPath, outputPath, protoMappingParams, cc);
 			});
 	}

@ -105,10 +95,7 @@ public class SparkBulkTagJob {

 		Dataset<String> datasources = readPath(
 			spark, inputPath
-				.substring(
-					0,
-					inputPath.lastIndexOf("/"))
-				+ "/datasource",
+				+ "datasource",
 			Datasource.class)
 				.filter((FilterFunction<Datasource>) ds -> isOKDatasource(ds))
 				.map((MapFunction<Datasource, String>) ds -> ds.getId(), Encoders.STRING());
@ -116,10 +103,10 @@ public class SparkBulkTagJob {
 		Map<String, List<Pair<String, SelectionConstraints>>> dsm = cc.getEoscDatasourceMap();

 		for (String ds : datasources.collectAsList()) {
-			final String dsId = ds.substring(3);
-			if (!dsm.containsKey(dsId)) {
+			// final String dsId = ds.substring(3);
+			if (!dsm.containsKey(ds)) {
 				ArrayList<Pair<String, SelectionConstraints>> eoscList = new ArrayList<>();
-				dsm.put(dsId, eoscList);
+				dsm.put(ds, eoscList);
 			}
 		}

@ -141,22 +128,30 @@ public class SparkBulkTagJob {
 		String inputPath,
 		String outputPath,
 		ProtoMap protoMappingParams,
-		Class<R> resultClazz,
 		CommunityConfiguration communityConfiguration) {

-		ResultTagger resultTagger = new ResultTagger();
-		readPath(spark, inputPath, resultClazz)
-			.map(patchResult(), Encoders.bean(resultClazz))
-			.filter(Objects::nonNull)
-			.map(
-				(MapFunction<R, R>) value -> resultTagger
-					.enrichContextCriteria(
-						value, communityConfiguration, protoMappingParams),
-				Encoders.bean(resultClazz))
-			.write()
-			.mode(SaveMode.Overwrite)
-			.option("compression", "gzip")
-			.json(outputPath);
+		ModelSupport.entityTypes
+			.keySet()
+			.parallelStream()
+			.filter(e -> ModelSupport.isResult(e))
+			.forEach(e -> {
+				removeOutputDir(spark, outputPath + e.name());
+				ResultTagger resultTagger = new ResultTagger();
+				Class<R> resultClazz = ModelSupport.entityTypes.get(e);
+				readPath(spark, inputPath + e.name(), resultClazz)
+					.map(patchResult(), Encoders.bean(resultClazz))
+					.filter(Objects::nonNull)
+					.map(
+						(MapFunction<R, R>) value -> resultTagger
+							.enrichContextCriteria(
+								value, communityConfiguration, protoMappingParams),
+						Encoders.bean(resultClazz))
+					.write()
+					.mode(SaveMode.Overwrite)
+					.option("compression", "gzip")
+					.json(outputPath + e.name());
+			});
+
 	}

 	public static <R> Dataset<R> readPath(
--- a/dhp-workflows/dhp-enrichment/src/main/java/eu/dnetlib/dhp/bulktag/community/Community.java
+++ b/dhp-workflows/dhp-enrichment/src/main/java/eu/dnetlib/dhp/bulktag/community/Community.java
@ -4,6 +4,7 @@ package eu.dnetlib.dhp.bulktag.community;
 import java.io.Serializable;
 import java.util.ArrayList;
 import java.util.List;
+import java.util.Optional;

 import com.google.gson.Gson;

@ -13,7 +14,7 @@ public class Community implements Serializable {
 	private String id;
 	private List<String> subjects = new ArrayList<>();
 	private List<Provider> providers = new ArrayList<>();
-	private List<ZenodoCommunity> zenodoCommunities = new ArrayList<>();
+	private List<String> zenodoCommunities = new ArrayList<>();
 	private SelectionConstraints constraints = new SelectionConstraints();
 	private SelectionConstraints removeConstraints = new SelectionConstraints();

@ -26,7 +27,7 @@ public class Community implements Serializable {
 		return !getSubjects().isEmpty()
 			|| !getProviders().isEmpty()
 			|| !getZenodoCommunities().isEmpty()
-			|| getConstraints().getCriteria() != null;
+			|| (Optional.ofNullable(getConstraints()).isPresent() && getConstraints().getCriteria() != null);
 	}

 	public String getId() {
@ -53,11 +54,11 @@ public class Community implements Serializable {
 		this.providers = providers;
 	}

-	public List<ZenodoCommunity> getZenodoCommunities() {
+	public List<String> getZenodoCommunities() {
 		return zenodoCommunities;
 	}

-	public void setZenodoCommunities(List<ZenodoCommunity> zenodoCommunities) {
+	public void setZenodoCommunities(List<String> zenodoCommunities) {
 		this.zenodoCommunities = zenodoCommunities;
 	}

--- a/dhp-workflows/dhp-enrichment/src/main/java/eu/dnetlib/dhp/bulktag/community/CommunityConfiguration.java
+++ b/dhp-workflows/dhp-enrichment/src/main/java/eu/dnetlib/dhp/bulktag/community/CommunityConfiguration.java
@ -81,7 +81,7 @@ public class CommunityConfiguration implements Serializable {
 		this.removeConstraintsMap = removeConstraintsMap;
 	}

-	CommunityConfiguration(final Map<String, Community> communities) {
+	public CommunityConfiguration(final Map<String, Community> communities) {
 		this.communities = communities;
 		init();
 	}
@ -117,10 +117,10 @@ public class CommunityConfiguration implements Serializable {
 				add(d.getOpenaireId(), new Pair<>(id, d.getSelectionConstraints()), datasourceMap);
 			}
 			// get zenodo communities
-			for (ZenodoCommunity zc : c.getZenodoCommunities()) {
+			for (String zc : c.getZenodoCommunities()) {
 				add(
-					zc.getZenodoCommunityId(),
-					new Pair<>(id, zc.getSelCriteria()),
+					zc,
+					new Pair<>(id, null),
 					zenodocommunityMap);
 			}
 			selectionConstraintsMap.put(id, c.getConstraints());
--- a/dhp-workflows/dhp-enrichment/src/main/java/eu/dnetlib/dhp/bulktag/community/CommunityConfigurationFactory.java
+++ b/dhp-workflows/dhp-enrichment/src/main/java/eu/dnetlib/dhp/bulktag/community/CommunityConfigurationFactory.java
@ -143,16 +143,16 @@ public class CommunityConfigurationFactory {
 		return providerList;
 	}

-	private static List<ZenodoCommunity> parseZenodoCommunities(final Node node) {
+	private static List<String> parseZenodoCommunities(final Node node) {

 		final List<Node> list = node.selectNodes("./zenodocommunities/zenodocommunity");
-		final List<ZenodoCommunity> zenodoCommunityList = new ArrayList<>();
+		final List<String> zenodoCommunityList = new ArrayList<>();
 		for (Node n : list) {
-			ZenodoCommunity zc = new ZenodoCommunity();
-			zc.setZenodoCommunityId(n.selectSingleNode("./zenodoid").getText());
-			zc.setSelCriteria(n.selectSingleNode("./selcriteria"));
+//			ZenodoCommunity zc = new ZenodoCommunity();
+//			zc.setZenodoCommunityId(n.selectSingleNode("./zenodoid").getText());
+//			zc.setSelCriteria(n.selectSingleNode("./selcriteria"));

-			zenodoCommunityList.add(zc);
+			zenodoCommunityList.add(n.selectSingleNode("./zenodoid").getText());
 		}

 		log.info("size of the zenodo community list " + zenodoCommunityList.size());
--- a/dhp-workflows/dhp-enrichment/src/main/java/eu/dnetlib/dhp/bulktag/community/Constraint.java
+++ b/dhp-workflows/dhp-enrichment/src/main/java/eu/dnetlib/dhp/bulktag/community/Constraint.java
@ -4,6 +4,8 @@ package eu.dnetlib.dhp.bulktag.community;
 import java.io.Serializable;
 import java.lang.reflect.InvocationTargetException;

+import org.apache.htrace.fasterxml.jackson.annotation.JsonIgnore;
+
 import eu.dnetlib.dhp.bulktag.criteria.Selection;
 import eu.dnetlib.dhp.bulktag.criteria.VerbResolver;

@ -12,6 +14,7 @@ public class Constraint implements Serializable {
 	private String field;
 	private String value;
 //	private String element;
+	@JsonIgnore
 	private Selection selection;

 	public String getVerb() {
@ -38,10 +41,11 @@ public class Constraint implements Serializable {
 		this.value = value;
 	}

-	public void setSelection(Selection sel) {
-		selection = sel;
-	}
-
+//@JsonIgnore
+	// public void setSelection(Selection sel) {
+//		selection = sel;
+//	}
+	@JsonIgnore
 	public void setSelection(VerbResolver resolver)
 		throws InvocationTargetException, NoSuchMethodException, InstantiationException,
 		IllegalAccessException {
@ -52,11 +56,4 @@ public class Constraint implements Serializable {
 		return selection.apply(metadata);
 	}

-//	public String getElement() {
-//		return element;
-//	}
-//
-//	public void setElement(String element) {
-//		this.element = element;
-//	}
 }
--- a/dhp-workflows/dhp-enrichment/src/main/java/eu/dnetlib/dhp/bulktag/community/QueryInformationSystem.java
+++ b/dhp-workflows/dhp-enrichment/src/main/java/eu/dnetlib/dhp/bulktag/community/QueryInformationSystem.java
@ -1,34 +0,0 @@
-
-package eu.dnetlib.dhp.bulktag.community;
-
-import java.io.IOException;
-import java.util.List;
-
-import org.apache.commons.io.IOUtils;
-import org.dom4j.DocumentException;
-import org.xml.sax.SAXException;
-
-import com.google.common.base.Joiner;
-
-import eu.dnetlib.dhp.utils.ISLookupClientFactory;
-import eu.dnetlib.enabling.is.lookup.rmi.ISLookUpException;
-import eu.dnetlib.enabling.is.lookup.rmi.ISLookUpService;
-
-public class QueryInformationSystem {
-
-	public static CommunityConfiguration getCommunityConfiguration(final String isLookupUrl)
-		throws ISLookUpException, DocumentException, SAXException, IOException {
-		ISLookUpService isLookUp = ISLookupClientFactory.getLookUpService(isLookupUrl);
-		final List<String> res = isLookUp
-			.quickSearchProfile(
-				IOUtils
-					.toString(
-						QueryInformationSystem.class
-							.getResourceAsStream(
-								"/eu/dnetlib/dhp/bulktag/query.xq")));
-
-		final String xmlConf = "<communities>" + Joiner.on(" ").join(res) + "</communities>";
-
-		return CommunityConfigurationFactory.newInstance(xmlConf);
-	}
-}
--- a/dhp-workflows/dhp-enrichment/src/main/java/eu/dnetlib/dhp/bulktag/community/ResultTagger.java
+++ b/dhp-workflows/dhp-enrichment/src/main/java/eu/dnetlib/dhp/bulktag/community/ResultTagger.java
@ -82,19 +82,23 @@ public class ResultTagger implements Serializable {
 		// communities contains all the communities to be not added to the context
 		final Set<String> removeCommunities = new HashSet<>();

+		// if (conf.getRemoveConstraintsMap().keySet().size() > 0)
 		conf
 			.getRemoveConstraintsMap()
 			.keySet()
-			.forEach(communityId -> {
-				if (conf.getRemoveConstraintsMap().get(communityId).getCriteria() != null &&
-					conf
-						.getRemoveConstraintsMap()
-						.get(communityId)
-						.getCriteria()
-						.stream()
-						.anyMatch(crit -> crit.verifyCriteria(param)))
-					removeCommunities.add(communityId);
-			});
+			.forEach(
+				communityId -> {
+					// log.info("Remove constraints for " + communityId);
+					if (conf.getRemoveConstraintsMap().keySet().contains(communityId) &&
+						conf.getRemoveConstraintsMap().get(communityId).getCriteria() != null &&
+						conf
+							.getRemoveConstraintsMap()
+							.get(communityId)
+							.getCriteria()
+							.stream()
+							.anyMatch(crit -> crit.verifyCriteria(param)))
+						removeCommunities.add(communityId);
+				});

 		// communities contains all the communities to be added as context for the result
 		final Set<String> communities = new HashSet<>();
@ -124,10 +128,10 @@ public class ResultTagger implements Serializable {
 		if (Objects.nonNull(result.getInstance())) {
 			for (Instance i : result.getInstance()) {
 				if (Objects.nonNull(i.getCollectedfrom()) && Objects.nonNull(i.getCollectedfrom().getKey())) {
-					collfrom.add(StringUtils.substringAfter(i.getCollectedfrom().getKey(), "|"));
+					collfrom.add(i.getCollectedfrom().getKey());
 				}
 				if (Objects.nonNull(i.getHostedby()) && Objects.nonNull(i.getHostedby().getKey())) {
-					hostdby.add(StringUtils.substringAfter(i.getHostedby().getKey(), "|"));
+					hostdby.add(i.getHostedby().getKey());
 				}

 			}
--- a/dhp-workflows/dhp-enrichment/src/main/java/eu/dnetlib/dhp/bulktag/community/SelectionConstraints.java
+++ b/dhp-workflows/dhp-enrichment/src/main/java/eu/dnetlib/dhp/bulktag/community/SelectionConstraints.java
@ -7,11 +7,13 @@ import java.util.Collection;
 import java.util.List;
 import java.util.Map;

+import com.fasterxml.jackson.annotation.JsonAutoDetect;
 import com.google.gson.Gson;
 import com.google.gson.reflect.TypeToken;

 import eu.dnetlib.dhp.bulktag.criteria.VerbResolver;

+@JsonAutoDetect
 public class SelectionConstraints implements Serializable {
 	private List<Constraints> criteria;

--- a/dhp-workflows/dhp-enrichment/src/main/java/eu/dnetlib/dhp/resulttocommunityfromorganization/PrepareResultCommunitySet.java
+++ b/dhp-workflows/dhp-enrichment/src/main/java/eu/dnetlib/dhp/resulttocommunityfromorganization/PrepareResultCommunitySet.java
@ -9,9 +9,7 @@ import java.util.*;
 import org.apache.commons.io.IOUtils;
 import org.apache.hadoop.io.compress.GzipCodec;
 import org.apache.spark.SparkConf;
-import org.apache.spark.api.java.function.FilterFunction;
 import org.apache.spark.api.java.function.MapFunction;
-import org.apache.spark.api.java.function.MapGroupsFunction;
 import org.apache.spark.sql.Dataset;
 import org.apache.spark.sql.Encoders;
 import org.apache.spark.sql.SparkSession;
@ -20,6 +18,8 @@ import org.slf4j.LoggerFactory;

 import com.google.gson.Gson;

+import eu.dnetlib.dhp.api.Utils;
+import eu.dnetlib.dhp.api.model.CommunityEntityMap;
 import eu.dnetlib.dhp.application.ArgumentApplicationParser;
 import eu.dnetlib.dhp.schema.common.ModelConstants;
 import eu.dnetlib.dhp.schema.oaf.Relation;
@ -48,10 +48,10 @@ public class PrepareResultCommunitySet {
 		final String outputPath = parser.get("outputPath");
 		log.info("outputPath: {}", outputPath);

-		final OrganizationMap organizationMap = new Gson()
-			.fromJson(
-				parser.get("organizationtoresultcommunitymap"),
-				OrganizationMap.class);
+		final boolean production = Boolean.valueOf(parser.get("production"));
+		log.info("production: {}", production);
+
+		final CommunityEntityMap organizationMap = Utils.getCommunityOrganization(production);
 		log.info("organizationMap: {}", new Gson().toJson(organizationMap));

 		SparkConf conf = new SparkConf();
@ -70,7 +70,7 @@ public class PrepareResultCommunitySet {
 		SparkSession spark,
 		String inputPath,
 		String outputPath,
-		OrganizationMap organizationMap) {
+		CommunityEntityMap organizationMap) {

 		Dataset<Relation> relation = readPath(spark, inputPath, Relation.class);
 		relation.createOrReplaceTempView("relation");
@ -115,7 +115,7 @@ public class PrepareResultCommunitySet {
 	}

 	private static MapFunction<ResultOrganizations, ResultCommunityList> mapResultCommunityFn(
-		OrganizationMap organizationMap) {
+		CommunityEntityMap organizationMap) {
 		return value -> {
 			String rId = value.getResultId();
 			Optional<List<String>> orgs = Optional.ofNullable(value.getMerges());
--- a/dhp-workflows/dhp-enrichment/src/main/java/eu/dnetlib/dhp/resulttocommunityfromorganization/SparkResultToCommunityFromOrganizationJob.java
+++ b/dhp-workflows/dhp-enrichment/src/main/java/eu/dnetlib/dhp/resulttocommunityfromorganization/SparkResultToCommunityFromOrganizationJob.java
@ -2,7 +2,7 @@
 package eu.dnetlib.dhp.resulttocommunityfromorganization;

 import static eu.dnetlib.dhp.PropagationConstant.*;
-import static eu.dnetlib.dhp.common.SparkSessionSupport.runWithSparkHiveSession;
+import static eu.dnetlib.dhp.common.SparkSessionSupport.runWithSparkSession;

 import java.util.ArrayList;
 import java.util.Arrays;
@ -22,6 +22,7 @@ import org.slf4j.LoggerFactory;

 import eu.dnetlib.dhp.application.ArgumentApplicationParser;
 import eu.dnetlib.dhp.schema.common.ModelConstants;
+import eu.dnetlib.dhp.schema.common.ModelSupport;
 import eu.dnetlib.dhp.schema.oaf.Context;
 import eu.dnetlib.dhp.schema.oaf.Result;
 import scala.Tuple2;
@ -53,29 +54,16 @@ public class SparkResultToCommunityFromOrganizationJob {
 		final String possibleupdatespath = parser.get("preparedInfoPath");
 		log.info("preparedInfoPath: {}", possibleupdatespath);

-		final String resultClassName = parser.get("resultTableName");
-		log.info("resultTableName: {}", resultClassName);
-
-		final Boolean saveGraph = Optional
-			.ofNullable(parser.get("saveGraph"))
-			.map(Boolean::valueOf)
-			.orElse(Boolean.TRUE);
-		log.info("saveGraph: {}", saveGraph);
-
-		@SuppressWarnings("unchecked")
-		Class<? extends Result> resultClazz = (Class<? extends Result>) Class.forName(resultClassName);
-
 		SparkConf conf = new SparkConf();
-		conf.set("hive.metastore.uris", parser.get("hive_metastore_uris"));

-		runWithSparkHiveSession(
+		runWithSparkSession(
 			conf,
 			isSparkSessionManaged,
 			spark -> {
-				removeOutputDir(spark, outputPath);
-				if (saveGraph) {
-					execPropagation(spark, inputPath, outputPath, resultClazz, possibleupdatespath);
-				}
+				// removeOutputDir(spark, outputPath);
+
+				execPropagation(spark, inputPath, outputPath, possibleupdatespath);
+
 			});
 	}

@ -83,22 +71,32 @@ public class SparkResultToCommunityFromOrganizationJob {
 		SparkSession spark,
 		String inputPath,
 		String outputPath,
-		Class<R> resultClazz,
 		String possibleUpdatesPath) {

 		Dataset<ResultCommunityList> possibleUpdates = readPath(spark, possibleUpdatesPath, ResultCommunityList.class);
-		Dataset<R> result = readPath(spark, inputPath, resultClazz);

-		result
-			.joinWith(
-				possibleUpdates,
-				result.col("id").equalTo(possibleUpdates.col("resultId")),
-				"left_outer")
-			.map(resultCommunityFn(), Encoders.bean(resultClazz))
-			.write()
-			.mode(SaveMode.Overwrite)
-			.option("compression", "gzip")
-			.json(outputPath);
+		ModelSupport.entityTypes
+			.keySet()
+			.parallelStream()
+			.forEach(e -> {
+				if (ModelSupport.isResult(e)) {
+					Class<R> resultClazz = ModelSupport.entityTypes.get(e);
+					removeOutputDir(spark, outputPath + e.name());
+					Dataset<R> result = readPath(spark, inputPath + e.name(), resultClazz);
+
+					result
+						.joinWith(
+							possibleUpdates,
+							result.col("id").equalTo(possibleUpdates.col("resultId")),
+							"left_outer")
+						.map(resultCommunityFn(), Encoders.bean(resultClazz))
+						.write()
+						.mode(SaveMode.Overwrite)
+						.option("compression", "gzip")
+						.json(outputPath + e.name());
+				}
+			});
+
 	}

 	private static <R extends Result> MapFunction<Tuple2<R, ResultCommunityList>, R> resultCommunityFn() {
--- a/dhp-workflows/dhp-enrichment/src/main/java/eu/dnetlib/dhp/resulttocommunityfromproject/PrepareResultCommunitySet.java
+++ b/dhp-workflows/dhp-enrichment/src/main/java/eu/dnetlib/dhp/resulttocommunityfromproject/PrepareResultCommunitySet.java
@ -0,0 +1,124 @@
+
+package eu.dnetlib.dhp.resulttocommunityfromproject;
+
+import static eu.dnetlib.dhp.PropagationConstant.*;
+import static eu.dnetlib.dhp.common.SparkSessionSupport.runWithSparkHiveSession;
+import static eu.dnetlib.dhp.common.SparkSessionSupport.runWithSparkSession;
+
+import java.util.*;
+
+import org.apache.commons.io.IOUtils;
+import org.apache.hadoop.io.compress.GzipCodec;
+import org.apache.spark.SparkConf;
+import org.apache.spark.api.java.function.MapFunction;
+import org.apache.spark.api.java.function.MapGroupsFunction;
+import org.apache.spark.sql.*;
+import org.apache.spark.sql.types.DataTypes;
+import org.apache.spark.sql.types.StructType;
+import org.slf4j.Logger;
+import org.slf4j.LoggerFactory;
+
+import com.google.gson.Gson;
+
+import eu.dnetlib.dhp.api.Utils;
+import eu.dnetlib.dhp.api.model.CommunityEntityMap;
+import eu.dnetlib.dhp.application.ArgumentApplicationParser;
+import eu.dnetlib.dhp.resulttocommunityfromorganization.ResultCommunityList;
+import eu.dnetlib.dhp.resulttocommunityfromorganization.ResultOrganizations;
+import eu.dnetlib.dhp.schema.common.ModelConstants;
+import eu.dnetlib.dhp.schema.oaf.Relation;
+import scala.Tuple2;
+
+public class PrepareResultCommunitySet {
+
+	private static final Logger log = LoggerFactory.getLogger(PrepareResultCommunitySet.class);
+
+	public static void main(String[] args) throws Exception {
+		String jsonConfiguration = IOUtils
+			.toString(
+				PrepareResultCommunitySet.class
+					.getResourceAsStream(
+						"/eu/dnetlib/dhp/resulttocommunityfromproject/input_preparecommunitytoresult_parameters.json"));
+
+		final ArgumentApplicationParser parser = new ArgumentApplicationParser(jsonConfiguration);
+		parser.parseArgument(args);
+
+		Boolean isSparkSessionManaged = isSparkSessionManaged(parser);
+		log.info("isSparkSessionManaged: {}", isSparkSessionManaged);
+
+		String inputPath = parser.get("sourcePath");
+		log.info("inputPath: {}", inputPath);
+
+		final String outputPath = parser.get("outputPath");
+		log.info("outputPath: {}", outputPath);
+
+		final boolean production = Boolean.valueOf(parser.get("production"));
+		log.info("production: {}", production);
+
+		final CommunityEntityMap projectsMap = Utils.getCommunityProjects(production);
+		log.info("projectsMap: {}", new Gson().toJson(projectsMap));
+
+		SparkConf conf = new SparkConf();
+
+		runWithSparkSession(
+			conf,
+			isSparkSessionManaged,
+			spark -> {
+				removeOutputDir(spark, outputPath);
+				prepareInfo(spark, inputPath, outputPath, projectsMap);
+			});
+	}
+
+	private static void prepareInfo(
+		SparkSession spark,
+		String inputPath,
+		String outputPath,
+		CommunityEntityMap projectMap) {
+
+		final StructType structureSchema = new StructType()
+			.add(
+				"dataInfo", new StructType()
+					.add("deletedbyinference", DataTypes.BooleanType)
+					.add("invisible", DataTypes.BooleanType))
+			.add("source", DataTypes.StringType)
+			.add("target", DataTypes.StringType)
+			.add("relClass", DataTypes.StringType);
+
+		spark
+			.read()
+			.schema(structureSchema)
+			.json(inputPath)
+			.filter(
+				"dataInfo.deletedbyinference != true " +
+					"and relClass == '" + ModelConstants.IS_PRODUCED_BY + "'")
+			.select(
+				new Column("source").as("resultId"),
+				new Column("target").as("projectId"))
+			.groupByKey((MapFunction<Row, String>) r -> (String) r.getAs("resultId"), Encoders.STRING())
+			.mapGroups((MapGroupsFunction<String, Row, ResultProjectList>) (k, v) -> {
+				ResultProjectList rpl = new ResultProjectList();
+				rpl.setResultId(k);
+				ArrayList<String> cl = new ArrayList<>();
+				cl.addAll(projectMap.get(v.next().getAs("projectId")));
+				v.forEachRemaining(r -> {
+					projectMap
+						.get(r.getAs("projectId"))
+						.forEach(c -> {
+							if (!cl.contains(c))
+								cl.add(c);
+						});
+
+				});
+				if (cl.size() == 0)
+					return null;
+				rpl.setCommunityList(cl);
+				return rpl;
+			}, Encoders.bean(ResultProjectList.class))
+			.filter(Objects::nonNull)
+			.write()
+			.mode(SaveMode.Overwrite)
+			.option("compression", "gzip")
+			.json(outputPath);
+	}
+
+}
--- a/dhp-workflows/dhp-enrichment/src/main/java/eu/dnetlib/dhp/resulttocommunityfromproject/ResultProjectList.java
+++ b/dhp-workflows/dhp-enrichment/src/main/java/eu/dnetlib/dhp/resulttocommunityfromproject/ResultProjectList.java
@ -0,0 +1,26 @@
+
+package eu.dnetlib.dhp.resulttocommunityfromproject;
+
+import java.io.Serializable;
+import java.util.ArrayList;
+
+public class ResultProjectList implements Serializable {
+	private String resultId;
+	private ArrayList<String> communityList;
+
+	public String getResultId() {
+		return resultId;
+	}
+
+	public void setResultId(String resultId) {
+		this.resultId = resultId;
+	}
+
+	public ArrayList<String> getCommunityList() {
+		return communityList;
+	}
+
+	public void setCommunityList(ArrayList<String> communityList) {
+		this.communityList = communityList;
+	}
+}
--- a/dhp-workflows/dhp-enrichment/src/main/java/eu/dnetlib/dhp/resulttocommunityfromproject/SparkResultToCommunityFromProject.java
+++ b/dhp-workflows/dhp-enrichment/src/main/java/eu/dnetlib/dhp/resulttocommunityfromproject/SparkResultToCommunityFromProject.java
@ -0,0 +1,149 @@
+
+package eu.dnetlib.dhp.resulttocommunityfromproject;
+
+import static eu.dnetlib.dhp.PropagationConstant.*;
+import static eu.dnetlib.dhp.PropagationConstant.PROPAGATION_RESULT_COMMUNITY_ORGANIZATION_CLASS_NAME;
+import static eu.dnetlib.dhp.common.SparkSessionSupport.runWithSparkHiveSession;
+import static eu.dnetlib.dhp.common.SparkSessionSupport.runWithSparkSession;
+
+import java.io.Serializable;
+import java.util.ArrayList;
+import java.util.Arrays;
+import java.util.List;
+import java.util.Optional;
+import java.util.stream.Collectors;
+
+import org.apache.commons.io.IOUtils;
+import org.apache.spark.SparkConf;
+import org.apache.spark.api.java.function.MapFunction;
+import org.apache.spark.sql.Dataset;
+import org.apache.spark.sql.Encoders;
+import org.apache.spark.sql.SaveMode;
+import org.apache.spark.sql.SparkSession;
+import org.slf4j.Logger;
+import org.slf4j.LoggerFactory;
+
+import eu.dnetlib.dhp.application.ArgumentApplicationParser;
+import eu.dnetlib.dhp.resulttocommunityfromorganization.ResultCommunityList;
+import eu.dnetlib.dhp.resulttocommunityfromorganization.SparkResultToCommunityFromOrganizationJob;
+import eu.dnetlib.dhp.schema.common.ModelConstants;
+import eu.dnetlib.dhp.schema.common.ModelSupport;
+import eu.dnetlib.dhp.schema.oaf.Context;
+import eu.dnetlib.dhp.schema.oaf.Result;
+import scala.Tuple2;
+
+/**
+ * @author miriam.baglioni
+ * @Date 11/10/23
+ */
+public class SparkResultToCommunityFromProject implements Serializable {
+	private static final Logger log = LoggerFactory.getLogger(SparkResultToCommunityFromProject.class);
+
+	public static void main(String[] args) throws Exception {
+		String jsonConfiguration = IOUtils
+			.toString(
+				SparkResultToCommunityFromProject.class
+					.getResourceAsStream(
+						"/eu/dnetlib/dhp/resulttocommunityfromproject/input_communitytoresult_parameters.json"));
+
+		final ArgumentApplicationParser parser = new ArgumentApplicationParser(jsonConfiguration);
+
+		parser.parseArgument(args);
+
+		Boolean isSparkSessionManaged = isSparkSessionManaged(parser);
+		log.info("isSparkSessionManaged: {}", isSparkSessionManaged);
+
+		String inputPath = parser.get("sourcePath");
+		log.info("inputPath: {}", inputPath);
+
+		final String outputPath = parser.get("outputPath");
+		log.info("outputPath: {}", outputPath);
+
+		final String possibleupdatespath = parser.get("preparedInfoPath");
+		log.info("preparedInfoPath: {}", possibleupdatespath);
+
+		SparkConf conf = new SparkConf();
+
+		runWithSparkSession(
+			conf,
+			isSparkSessionManaged,
+			spark -> {
+
+				execPropagation(spark, inputPath, outputPath, possibleupdatespath);
+
+			});
+	}
+
+	private static <R extends Result> void execPropagation(
+		SparkSession spark,
+		String inputPath,
+		String outputPath,
+
+		String possibleUpdatesPath) {
+
+		Dataset<ResultProjectList> possibleUpdates = readPath(spark, possibleUpdatesPath, ResultProjectList.class);
+
+		ModelSupport.entityTypes
+			.keySet()
+			.parallelStream()
+			.forEach(e -> {
+				if (ModelSupport.isResult(e)) {
+					removeOutputDir(spark, outputPath + e.name());
+					Class<R> resultClazz = ModelSupport.entityTypes.get(e);
+					Dataset<R> result = readPath(spark, inputPath + e.name(), resultClazz);
+
+					result
+						.joinWith(
+							possibleUpdates,
+							result.col("id").equalTo(possibleUpdates.col("resultId")),
+							"left_outer")
+						.map(resultCommunityFn(), Encoders.bean(resultClazz))
+						.write()
+						.mode(SaveMode.Overwrite)
+						.option("compression", "gzip")
+						.json(outputPath + e.name());
+				}
+			});
+
+	}
+
+	private static <R extends Result> MapFunction<Tuple2<R, ResultProjectList>, R> resultCommunityFn() {
+		return value -> {
+			R ret = value._1();
+			Optional<ResultProjectList> rcl = Optional.ofNullable(value._2());
+			if (rcl.isPresent()) {
+				ArrayList<String> communitySet = rcl.get().getCommunityList();
+				List<String> contextList = ret
+					.getContext()
+					.stream()
+					.map(Context::getId)
+					.collect(Collectors.toList());
+
+				@SuppressWarnings("unchecked")
+				R res = (R) ret.getClass().newInstance();
+
+				res.setId(ret.getId());
+				List<Context> propagatedContexts = new ArrayList<>();
+				for (String cId : communitySet) {
+					if (!contextList.contains(cId)) {
+						Context newContext = new Context();
+						newContext.setId(cId);
+						newContext
+							.setDataInfo(
+								Arrays
+									.asList(
+										getDataInfo(
+											PROPAGATION_DATA_INFO_TYPE,
+											PROPAGATION_RESULT_COMMUNITY_PROJECT_CLASS_ID,
+											PROPAGATION_RESULT_COMMUNITY_PROJECT_CLASS_NAME,
+											ModelConstants.DNET_PROVENANCE_ACTIONS)));
+						propagatedContexts.add(newContext);
+					}
+				}
+				res.setContext(propagatedContexts);
+				ret.mergeFrom(res);
+			}
+			return ret;
+		};
+	}
+}
--- a/dhp-workflows/dhp-enrichment/src/main/resources/eu/dnetlib/dhp/bulktag/input_bulkTag_parameters.json
+++ b/dhp-workflows/dhp-enrichment/src/main/resources/eu/dnetlib/dhp/bulktag/input_bulkTag_parameters.json
@ -1,10 +1,5 @@
 [
-  {
-    "paramName":"is",
-    "paramLongName":"isLookUpUrl",
-    "paramDescription": "URL of the isLookUp Service",
-    "paramRequired": true
-  },
+
  {
    "paramName":"s",
    "paramLongName":"sourcePath",
@ -17,12 +12,7 @@
    "paramDescription": "the json path associated to each selection field",
    "paramRequired": true
  },
-  {
-    "paramName":"tn",
-    "paramLongName":"resultTableName",
-    "paramDescription": "the name of the result table we are currently working on",
-    "paramRequired": true
-  },
+
  {
    "paramName": "out",
    "paramLongName": "outputPath",
@ -35,17 +25,19 @@
    "paramDescription": "true if the spark session is managed, false otherwise",
    "paramRequired": false
  },
-  {
-    "paramName": "test",
-    "paramLongName": "isTest",
-    "paramDescription": "Parameter intended for testing purposes only. True if the reun is relatesd to a test and so the taggingConf parameter should be loaded",
-    "paramRequired": false
-  },
+
  {
    "paramName": "tg",
    "paramLongName": "taggingConf",
    "paramDescription": "this parameter is intended for testing purposes only. It is a possible tagging configuration obtained via the XQUERY. Intended to be removed",
    "paramRequired": false
+  },
+
+  {
+    "paramName": "p",
+    "paramLongName": "production",
+    "paramDescription": "this parameter is intended for testing purposes only. It is a possible tagging configuration obtained via the XQUERY. Intended to be removed",
+    "paramRequired": true
  }

 ]
--- a/dhp-workflows/dhp-enrichment/src/main/resources/eu/dnetlib/dhp/bulktag/oozie_app/workflow.xml
+++ b/dhp-workflows/dhp-enrichment/src/main/resources/eu/dnetlib/dhp/bulktag/oozie_app/workflow.xml
@ -4,10 +4,6 @@
            <name>sourcePath</name>
            <description>the source path</description>
        </property>
-        <property>
-            <name>isLookUpUrl</name>
-            <description>the isLookup service endpoint</description>
-        </property>
        <property>
            <name>pathMap</name>
            <description>the json path associated to each selection field</description>
@ -44,7 +40,7 @@
        </configuration>
    </global>

-    <start to="reset_outputpath"/>
+    <start to="exec_bulktag"/>

    <kill name="Kill">
        <message>Action failed, error message[${wf:errorMessage(wf:lastErrorNode())}]</message>
@ -102,16 +98,9 @@
        <error to="Kill"/>
    </action>

-    <join name="copy_wait" to="fork_exec_bulktag"/>
+    <join name="copy_wait" to="exec_bulktag"/>

-    <fork name="fork_exec_bulktag">
-        <path start="bulktag_publication"/>
-        <path start="bulktag_dataset"/>
-        <path start="bulktag_otherresearchproduct"/>
-        <path start="bulktag_software"/>
-    </fork>
-
-    <action name="bulktag_publication">
+    <action name="exec_bulktag">
        <spark xmlns="uri:oozie:spark-action:0.2">
            <master>yarn-cluster</master>
            <mode>cluster</mode>
@ -128,98 +117,15 @@
                --conf spark.yarn.historyServer.address=${spark2YarnHistoryServerAddress}
                --conf spark.eventLog.dir=${nameNode}${spark2EventLogDir}
            </spark-opts>
-            <arg>--sourcePath</arg><arg>${sourcePath}/publication</arg>
-            <arg>--resultTableName</arg><arg>eu.dnetlib.dhp.schema.oaf.Publication</arg>
-            <arg>--outputPath</arg><arg>${outputPath}/publication</arg>
+            <arg>--sourcePath</arg><arg>${sourcePath}/</arg>
+            <arg>--outputPath</arg><arg>${outputPath}/</arg>
            <arg>--pathMap</arg><arg>${pathMap}</arg>
-            <arg>--isLookUpUrl</arg><arg>${isLookUpUrl}</arg>
+            <arg>--production</arg><arg>${production}</arg>
        </spark>
-        <ok to="wait"/>
+        <ok to="End"/>
        <error to="Kill"/>
    </action>

-    <action name="bulktag_dataset">
-        <spark xmlns="uri:oozie:spark-action:0.2">
-            <master>yarn-cluster</master>
-            <mode>cluster</mode>
-            <name>bulkTagging-dataset</name>
-            <class>eu.dnetlib.dhp.bulktag.SparkBulkTagJob</class>
-            <jar>dhp-enrichment-${projectVersion}.jar</jar>
-            <spark-opts>
-                --num-executors=${sparkExecutorNumber}
-                --executor-memory=${sparkExecutorMemory}
-                --executor-cores=${sparkExecutorCores}
-                --driver-memory=${sparkDriverMemory}
-                --conf spark.extraListeners=${spark2ExtraListeners}
-                --conf spark.sql.queryExecutionListeners=${spark2SqlQueryExecutionListeners}
-                --conf spark.yarn.historyServer.address=${spark2YarnHistoryServerAddress}
-                --conf spark.eventLog.dir=${nameNode}${spark2EventLogDir}
-            </spark-opts>
-            <arg>--sourcePath</arg><arg>${sourcePath}/dataset</arg>
-            <arg>--resultTableName</arg><arg>eu.dnetlib.dhp.schema.oaf.Dataset</arg>
-            <arg>--outputPath</arg><arg>${outputPath}/dataset</arg>
-            <arg>--pathMap</arg><arg>${pathMap}</arg>
-            <arg>--isLookUpUrl</arg><arg>${isLookUpUrl}</arg>
-        </spark>
-        <ok to="wait"/>
-        <error to="Kill"/>
-    </action>
-
-    <action name="bulktag_otherresearchproduct">
-        <spark xmlns="uri:oozie:spark-action:0.2">
-            <master>yarn-cluster</master>
-            <mode>cluster</mode>
-            <name>bulkTagging-orp</name>
-            <class>eu.dnetlib.dhp.bulktag.SparkBulkTagJob</class>
-            <jar>dhp-enrichment-${projectVersion}.jar</jar>
-            <spark-opts>
-                --num-executors=${sparkExecutorNumber}
-                --executor-memory=${sparkExecutorMemory}
-                --executor-cores=${sparkExecutorCores}
-                --driver-memory=${sparkDriverMemory}
-                --conf spark.extraListeners=${spark2ExtraListeners}
-                --conf spark.sql.queryExecutionListeners=${spark2SqlQueryExecutionListeners}
-                --conf spark.yarn.historyServer.address=${spark2YarnHistoryServerAddress}
-                --conf spark.eventLog.dir=${nameNode}${spark2EventLogDir}
-            </spark-opts>
-            <arg>--sourcePath</arg><arg>${sourcePath}/otherresearchproduct</arg>
-            <arg>--resultTableName</arg><arg>eu.dnetlib.dhp.schema.oaf.OtherResearchProduct</arg>
-            <arg>--outputPath</arg><arg>${outputPath}/otherresearchproduct</arg>
-            <arg>--pathMap</arg><arg>${pathMap}</arg>
-            <arg>--isLookUpUrl</arg><arg>${isLookUpUrl}</arg>
-        </spark>
-        <ok to="wait"/>
-        <error to="Kill"/>
-    </action>
-
-    <action name="bulktag_software">
-        <spark xmlns="uri:oozie:spark-action:0.2">
-            <master>yarn-cluster</master>
-            <mode>cluster</mode>
-            <name>bulkTagging-software</name>
-            <class>eu.dnetlib.dhp.bulktag.SparkBulkTagJob</class>
-            <jar>dhp-enrichment-${projectVersion}.jar</jar>
-            <spark-opts>
-                --num-executors=${sparkExecutorNumber}
-                --executor-memory=${sparkExecutorMemory}
-                --executor-cores=${sparkExecutorCores}
-                --driver-memory=${sparkDriverMemory}
-                --conf spark.extraListeners=${spark2ExtraListeners}
-                --conf spark.sql.queryExecutionListeners=${spark2SqlQueryExecutionListeners}
-                --conf spark.yarn.historyServer.address=${spark2YarnHistoryServerAddress}
-                --conf spark.eventLog.dir=${nameNode}${spark2EventLogDir}
-            </spark-opts>
-            <arg>--sourcePath</arg><arg>${sourcePath}/software</arg>
-            <arg>--resultTableName</arg><arg>eu.dnetlib.dhp.schema.oaf.Software</arg>
-            <arg>--outputPath</arg><arg>${outputPath}/software</arg>
-            <arg>--pathMap</arg><arg>${pathMap}</arg>
-            <arg>--isLookUpUrl</arg><arg>${isLookUpUrl}</arg>
-        </spark>
-        <ok to="wait"/>
-        <error to="Kill"/>
-    </action>
-
-    <join name="wait" to="End"/>


    <end name="End"/>
--- a/dhp-workflows/dhp-enrichment/src/main/resources/eu/dnetlib/dhp/bulktag/query.xq
+++ b/dhp-workflows/dhp-enrichment/src/main/resources/eu/dnetlib/dhp/bulktag/query.xq
@ -1,62 +0,0 @@
-for $x in collection('/db/DRIVER/ContextDSResources/ContextDSResourceType')
-let $subj := $x//CONFIGURATION/context/param[./@name='subject']/text()
-let $datasources := $x//CONFIGURATION/context/category[./@id=concat($x//CONFIGURATION/context/@id,'::contentproviders')]/concept
-let $organizations := $x//CONFIGURATION/context/category[./@id=concat($x//CONFIGURATION/context/@id,'::resultorganizations')]/concept
-let $communities := $x//CONFIGURATION/context/category[./@id=concat($x//CONFIGURATION/context/@id,'::zenodocommunities')]/concept
-let $fos := $x//CONFIGURATION/context/param[./@name='fos']/text()
-let $sdg := $x//CONFIGURATION/context/param[./@name='sdg']/text()
-let $zenodo := $x//param[./@name='zenodoCommunity']/text()
-where $x//CONFIGURATION/context[./@type='community' or ./@type='ri'] and $x//context/param[./@name = 'status']/text() != 'hidden'
-return
-<community>
-{ $x//CONFIGURATION/context/@id}
-<removeConstraints>
-{$x//CONFIGURATION/context/param[./@name='removeConstraints']/text() }
-</removeConstraints>
-<advancedConstraints>
-{$x//CONFIGURATION/context/param[./@name='advancedConstraints']/text() }
-</advancedConstraints>
- <subjects>
- {for $y in tokenize($subj,',')
- return
- <subject>{$y}</subject>}
- {for $y in tokenize($fos,',')
- return
- <subject>{$y}</subject>}
- {for $y in tokenize($sdg,',')
- return
- <subject>{$y}</subject>}
- </subjects>
- <datasources>
- {for $d in $datasources
- where $d/param[./@name='enabled']/text()='true'
- return
- <datasource>
- <openaireId>
- {$d//param[./@name='openaireId']/text()}
- </openaireId>
- <selcriteria>
- {$d/param[./@name='selcriteria']/text()}
- </selcriteria>
- </datasource> }
- </datasources>
-  <zenodocommunities>
-{for $zc in $zenodo
-return
-<zenodocommunity>
-<zenodoid>
-{$zc}
-</zenodoid>
-</zenodocommunity>}
-{for $zc in $communities
-return
-<zenodocommunity>
-<zenodoid>
-{$zc/param[./@name='zenodoid']/text()}
-</zenodoid>
-<selcriteria>
-{$zc/param[./@name='selcriteria']/text()}
-</selcriteria>
-</zenodocommunity>}
-</zenodocommunities>
-</community>
--- a/dhp-workflows/dhp-enrichment/src/main/resources/eu/dnetlib/dhp/resulttocommunityfromorganization/input_communitytoresult_parameters.json
+++ b/dhp-workflows/dhp-enrichment/src/main/resources/eu/dnetlib/dhp/resulttocommunityfromorganization/input_communitytoresult_parameters.json
@ -5,24 +5,7 @@
    "paramDescription": "the path of the sequencial file to read",
    "paramRequired": true
  },
-  {
-    "paramName":"h",
-    "paramLongName":"hive_metastore_uris",
-    "paramDescription": "the hive metastore uris",
-    "paramRequired": true
-  },
-  {
-    "paramName":"sg",
-    "paramLongName":"saveGraph",
-    "paramDescription": "true if the new version of the graph must be saved",
-    "paramRequired": false
-  },
-  {
-    "paramName":"test",
-    "paramLongName":"isTest",
-    "paramDescription": "true if it is executing a test",
-    "paramRequired": false
-  },
+
  {
    "paramName": "out",
    "paramLongName": "outputPath",
@ -35,12 +18,6 @@
    "paramDescription": "true if the spark session is managed, false otherwise",
    "paramRequired": false
  },
-  {
-    "paramName":"tn",
-    "paramLongName":"resultTableName",
-    "paramDescription": "the name of the result table we are currently working on",
-    "paramRequired": true
-  },
  {
    "paramName": "p",
    "paramLongName": "preparedInfoPath",
--- a/dhp-workflows/dhp-enrichment/src/main/resources/eu/dnetlib/dhp/resulttocommunityfromorganization/input_preparecommunitytoresult_parameters.json
+++ b/dhp-workflows/dhp-enrichment/src/main/resources/eu/dnetlib/dhp/resulttocommunityfromorganization/input_preparecommunitytoresult_parameters.json
@ -5,12 +5,6 @@
    "paramDescription": "the path of the sequencial file to read",
    "paramRequired": true
  },
-  {
-    "paramName":"ocm",
-    "paramLongName":"organizationtoresultcommunitymap",
-    "paramDescription": "the map for the association organization communities",
-    "paramRequired": true
-  },
  {
    "paramName":"h",
    "paramLongName":"hive_metastore_uris",
@ -28,6 +22,12 @@
    "paramLongName": "outputPath",
    "paramDescription": "the path used to store temporary output files",
    "paramRequired": true
+  },
+  {
+    "paramName": "p",
+    "paramLongName": "production",
+    "paramDescription": "the path used to store temporary output files",
+    "paramRequired": true
  }

 ]
--- a/dhp-workflows/dhp-enrichment/src/main/resources/eu/dnetlib/dhp/resulttocommunityfromorganization/oozie_app/workflow.xml
+++ b/dhp-workflows/dhp-enrichment/src/main/resources/eu/dnetlib/dhp/resulttocommunityfromorganization/oozie_app/workflow.xml
@ -4,10 +4,7 @@
            <name>sourcePath</name>
            <description>the source path</description>
        </property>
-        <property>
-            <name>organizationtoresultcommunitymap</name>
-            <description>organization community map</description>
-        </property>
+
        <property>
            <name>outputPath</name>
            <description>the output path</description>
@ -25,7 +22,7 @@
        </configuration>
    </global>

-    <start to="reset_outputpath"/>
+    <start to="prepare_result_communitylist"/>

    <kill name="Kill">
        <message>Action failed, error message[${wf:errorMessage(wf:lastErrorNode())}]</message>
@ -93,33 +90,28 @@
            <class>eu.dnetlib.dhp.resulttocommunityfromorganization.PrepareResultCommunitySet</class>
            <jar>dhp-enrichment-${projectVersion}.jar</jar>
            <spark-opts>
-                --executor-cores=${sparkExecutorCores}
-                --executor-memory=${sparkExecutorMemory}
+                --executor-cores=6
+                --executor-memory=5G
+                --conf spark.executor.memoryOverhead=3g
+                --conf spark.sql.shuffle.partitions=3284
                --driver-memory=${sparkDriverMemory}
                --conf spark.extraListeners=${spark2ExtraListeners}
                --conf spark.sql.queryExecutionListeners=${spark2SqlQueryExecutionListeners}
                --conf spark.yarn.historyServer.address=${spark2YarnHistoryServerAddress}
                --conf spark.eventLog.dir=${nameNode}${spark2EventLogDir}
-                --conf spark.dynamicAllocation.enabled=true
+
                --conf spark.dynamicAllocation.maxExecutors=${spark2MaxExecutors}
            </spark-opts>
            <arg>--sourcePath</arg><arg>${sourcePath}/relation</arg>
            <arg>--outputPath</arg><arg>${workingDir}/preparedInfo/resultCommunityList</arg>
            <arg>--hive_metastore_uris</arg><arg>${hive_metastore_uris}</arg>
-            <arg>--organizationtoresultcommunitymap</arg><arg>${organizationtoresultcommunitymap}</arg>
+            <arg>--production</arg><arg>${production}</arg>
        </spark>
-        <ok to="fork-join-exec-propagation"/>
+        <ok to="exec-propagation"/>
        <error to="Kill"/>
    </action>

-    <fork name="fork-join-exec-propagation">
-        <path start="join_propagate_publication"/>
-        <path start="join_propagate_dataset"/>
-        <path start="join_propagate_otherresearchproduct"/>
-        <path start="join_propagate_software"/>
-    </fork>
-
-    <action name="join_propagate_publication">
+    <action name="exec-propagation">
        <spark xmlns="uri:oozie:spark-action:0.2">
            <master>yarn</master>
            <mode>cluster</mode>
@ -127,115 +119,26 @@
            <class>eu.dnetlib.dhp.resulttocommunityfromorganization.SparkResultToCommunityFromOrganizationJob</class>
            <jar>dhp-enrichment-${projectVersion}.jar</jar>
            <spark-opts>
-                --executor-cores=${sparkExecutorCores}
-                --executor-memory=${sparkExecutorMemory}
+                --executor-cores=6
+                --executor-memory=5G
+                --conf spark.executor.memoryOverhead=3g
+                --conf spark.sql.shuffle.partitions=3284
                --driver-memory=${sparkDriverMemory}
                --conf spark.extraListeners=${spark2ExtraListeners}
                --conf spark.sql.queryExecutionListeners=${spark2SqlQueryExecutionListeners}
                --conf spark.yarn.historyServer.address=${spark2YarnHistoryServerAddress}
                --conf spark.eventLog.dir=${nameNode}${spark2EventLogDir}
-                --conf spark.dynamicAllocation.enabled=true
                --conf spark.dynamicAllocation.maxExecutors=${spark2MaxExecutors}
            </spark-opts>
            <arg>--preparedInfoPath</arg><arg>${workingDir}/preparedInfo/resultCommunityList</arg>
-            <arg>--sourcePath</arg><arg>${sourcePath}/publication</arg>
-            <arg>--outputPath</arg><arg>${outputPath}/publication</arg>
-            <arg>--hive_metastore_uris</arg><arg>${hive_metastore_uris}</arg>
-            <arg>--resultTableName</arg><arg>eu.dnetlib.dhp.schema.oaf.Publication</arg>
-            <arg>--saveGraph</arg><arg>${saveGraph}</arg>
+            <arg>--sourcePath</arg><arg>${sourcePath}/</arg>
+            <arg>--outputPath</arg><arg>${outputPath}/</arg>
        </spark>
-        <ok to="wait2"/>
+        <ok to="End"/>
        <error to="Kill"/>
    </action>

-    <action name="join_propagate_dataset">
-        <spark xmlns="uri:oozie:spark-action:0.2">
-            <master>yarn</master>
-            <mode>cluster</mode>
-            <name>community2resultfromorganization-Dataset</name>
-            <class>eu.dnetlib.dhp.resulttocommunityfromorganization.SparkResultToCommunityFromOrganizationJob</class>
-            <jar>dhp-enrichment-${projectVersion}.jar</jar>
-            <spark-opts>
-                --executor-cores=${sparkExecutorCores}
-                --executor-memory=${sparkExecutorMemory}
-                --driver-memory=${sparkDriverMemory}
-                --conf spark.extraListeners=${spark2ExtraListeners}
-                --conf spark.sql.queryExecutionListeners=${spark2SqlQueryExecutionListeners}
-                --conf spark.yarn.historyServer.address=${spark2YarnHistoryServerAddress}
-                --conf spark.eventLog.dir=${nameNode}${spark2EventLogDir}
-                --conf spark.dynamicAllocation.enabled=true
-                --conf spark.dynamicAllocation.maxExecutors=${spark2MaxExecutors}
-            </spark-opts>
-            <arg>--preparedInfoPath</arg><arg>${workingDir}/preparedInfo/resultCommunityList</arg>
-            <arg>--sourcePath</arg><arg>${sourcePath}/dataset</arg>
-            <arg>--outputPath</arg><arg>${outputPath}/dataset</arg>
-            <arg>--hive_metastore_uris</arg><arg>${hive_metastore_uris}</arg>
-            <arg>--resultTableName</arg><arg>eu.dnetlib.dhp.schema.oaf.Dataset</arg>
-            <arg>--saveGraph</arg><arg>${saveGraph}</arg>
-        </spark>
-        <ok to="wait2"/>
-        <error to="Kill"/>
-    </action>

-    <action name="join_propagate_otherresearchproduct">
-        <spark xmlns="uri:oozie:spark-action:0.2">
-            <master>yarn</master>
-            <mode>cluster</mode>
-            <name>community2resultfromorganization-ORP</name>
-            <class>eu.dnetlib.dhp.resulttocommunityfromorganization.SparkResultToCommunityFromOrganizationJob</class>
-            <jar>dhp-enrichment-${projectVersion}.jar</jar>
-            <spark-opts>
-                --executor-cores=${sparkExecutorCores}
-                --executor-memory=${sparkExecutorMemory}
-                --driver-memory=${sparkDriverMemory}
-                --conf spark.extraListeners=${spark2ExtraListeners}
-                --conf spark.sql.queryExecutionListeners=${spark2SqlQueryExecutionListeners}
-                --conf spark.yarn.historyServer.address=${spark2YarnHistoryServerAddress}
-                --conf spark.eventLog.dir=${nameNode}${spark2EventLogDir}
-                --conf spark.dynamicAllocation.enabled=true
-                --conf spark.dynamicAllocation.maxExecutors=${spark2MaxExecutors}
-            </spark-opts>
-            <arg>--preparedInfoPath</arg><arg>${workingDir}/preparedInfo/resultCommunityList</arg>
-            <arg>--sourcePath</arg><arg>${sourcePath}/otherresearchproduct</arg>
-            <arg>--outputPath</arg><arg>${outputPath}/otherresearchproduct</arg>
-            <arg>--hive_metastore_uris</arg><arg>${hive_metastore_uris}</arg>
-            <arg>--resultTableName</arg><arg>eu.dnetlib.dhp.schema.oaf.OtherResearchProduct</arg>
-            <arg>--saveGraph</arg><arg>${saveGraph}</arg>
-        </spark>
-        <ok to="wait2"/>
-        <error to="Kill"/>
-    </action>
-
-    <action name="join_propagate_software">
-        <spark xmlns="uri:oozie:spark-action:0.2">
-            <master>yarn</master>
-            <mode>cluster</mode>
-            <name>community2resultfromorganization-Software</name>
-            <class>eu.dnetlib.dhp.resulttocommunityfromorganization.SparkResultToCommunityFromOrganizationJob</class>
-            <jar>dhp-enrichment-${projectVersion}.jar</jar>
-            <spark-opts>
-                --executor-cores=${sparkExecutorCores}
-                --executor-memory=${sparkExecutorMemory}
-                --driver-memory=${sparkDriverMemory}
-                --conf spark.extraListeners=${spark2ExtraListeners}
-                --conf spark.sql.queryExecutionListeners=${spark2SqlQueryExecutionListeners}
-                --conf spark.yarn.historyServer.address=${spark2YarnHistoryServerAddress}
-                --conf spark.eventLog.dir=${nameNode}${spark2EventLogDir}
-                --conf spark.dynamicAllocation.enabled=true
-                --conf spark.dynamicAllocation.maxExecutors=${spark2MaxExecutors}
-            </spark-opts>
-            <arg>--preparedInfoPath</arg><arg>${workingDir}/preparedInfo/resultCommunityList</arg>
-            <arg>--sourcePath</arg><arg>${sourcePath}/software</arg>
-            <arg>--outputPath</arg><arg>${outputPath}/software</arg>
-            <arg>--hive_metastore_uris</arg><arg>${hive_metastore_uris}</arg>
-            <arg>--resultTableName</arg><arg>eu.dnetlib.dhp.schema.oaf.Software</arg>
-            <arg>--saveGraph</arg><arg>${saveGraph}</arg>
-        </spark>
-        <ok to="wait2"/>
-        <error to="Kill"/>
-    </action>
-
-    <join name="wait2" to="End"/>

    <end name="End"/>

--- a/dhp-workflows/dhp-enrichment/src/main/resources/eu/dnetlib/dhp/resulttocommunityfromproject/input_communitytoresult_parameters.json
+++ b/dhp-workflows/dhp-enrichment/src/main/resources/eu/dnetlib/dhp/resulttocommunityfromproject/input_communitytoresult_parameters.json
@ -0,0 +1,28 @@
+[
+  {
+    "paramName":"s",
+    "paramLongName":"sourcePath",
+    "paramDescription": "the path of the sequencial file to read",
+    "paramRequired": true
+  },
+
+  {
+    "paramName": "out",
+    "paramLongName": "outputPath",
+    "paramDescription": "the path used to store temporary output files",
+    "paramRequired": true
+  },
+  {
+    "paramName": "ssm",
+    "paramLongName": "isSparkSessionManaged",
+    "paramDescription": "true if the spark session is managed, false otherwise",
+    "paramRequired": false
+  },
+  {
+    "paramName": "p",
+    "paramLongName": "preparedInfoPath",
+    "paramDescription": "the path where prepared info have been stored",
+    "paramRequired": true
+  }
+
+]
--- a/dhp-workflows/dhp-enrichment/src/main/resources/eu/dnetlib/dhp/resulttocommunityfromproject/input_preparecommunitytoresult_parameters.json
+++ b/dhp-workflows/dhp-enrichment/src/main/resources/eu/dnetlib/dhp/resulttocommunityfromproject/input_preparecommunitytoresult_parameters.json
@ -0,0 +1,28 @@
+[
+  {
+    "paramName":"s",
+    "paramLongName":"sourcePath",
+    "paramDescription": "the path of the sequencial file to read",
+    "paramRequired": true
+  },
+
+  {
+    "paramName": "ssm",
+    "paramLongName": "isSparkSessionManaged",
+    "paramDescription": "true if the spark session is managed, false otherwise",
+    "paramRequired": false
+  },
+  {
+    "paramName": "out",
+    "paramLongName": "outputPath",
+    "paramDescription": "the path used to store temporary output files",
+    "paramRequired": true
+  },
+  {
+    "paramName": "p",
+    "paramLongName": "production",
+    "paramDescription": "the path used to store temporary output files",
+    "paramRequired": true
+  }
+
+]
--- a/dhp-workflows/dhp-enrichment/src/main/resources/eu/dnetlib/dhp/resulttocommunityfromproject/oozie_app/config-default.xml
+++ b/dhp-workflows/dhp-enrichment/src/main/resources/eu/dnetlib/dhp/resulttocommunityfromproject/oozie_app/config-default.xml
@ -0,0 +1,58 @@
+<configuration>
+    <property>
+        <name>jobTracker</name>
+        <value>yarnRM</value>
+    </property>
+    <property>
+        <name>nameNode</name>
+        <value>hdfs://nameservice1</value>
+    </property>
+    <property>
+        <name>oozie.use.system.libpath</name>
+        <value>true</value>
+    </property>
+    <property>
+        <name>oozie.action.sharelib.for.spark</name>
+        <value>spark2</value>
+    </property>
+    <property>
+        <name>hive_metastore_uris</name>
+        <value>thrift://iis-cdh5-test-m3.ocean.icm.edu.pl:9083</value>
+    </property>
+    <property>
+        <name>spark2YarnHistoryServerAddress</name>
+        <value>http://iis-cdh5-test-gw.ocean.icm.edu.pl:18089</value>
+    </property>
+    <property>
+        <name>spark2EventLogDir</name>
+        <value>/user/spark/spark2ApplicationHistory</value>
+    </property>
+    <property>
+        <name>spark2ExtraListeners</name>
+        <value>com.cloudera.spark.lineage.NavigatorAppListener</value>
+    </property>
+    <property>
+        <name>spark2SqlQueryExecutionListeners</name>
+        <value>com.cloudera.spark.lineage.NavigatorQueryListener</value>
+    </property>
+    <property>
+        <name>sparkExecutorNumber</name>
+        <value>4</value>
+    </property>
+    <property>
+        <name>sparkDriverMemory</name>
+        <value>15G</value>
+    </property>
+    <property>
+        <name>sparkExecutorMemory</name>
+        <value>6G</value>
+    </property>
+    <property>
+        <name>sparkExecutorCores</name>
+        <value>1</value>
+    </property>
+    <property>
+        <name>spark2MaxExecutors</name>
+        <value>50</value>
+    </property>
+</configuration>
--- a/dhp-workflows/dhp-enrichment/src/main/resources/eu/dnetlib/dhp/resulttocommunityfromproject/oozie_app/workflow.xml
+++ b/dhp-workflows/dhp-enrichment/src/main/resources/eu/dnetlib/dhp/resulttocommunityfromproject/oozie_app/workflow.xml
@ -0,0 +1,144 @@
+<workflow-app name="community_to_result_propagation_project" xmlns="uri:oozie:workflow:0.5">
+    <parameters>
+        <property>
+            <name>sourcePath</name>
+            <description>the source path</description>
+        </property>
+
+        <property>
+            <name>outputPath</name>
+            <description>the output path</description>
+        </property>
+    </parameters>
+
+    <global>
+        <job-tracker>${jobTracker}</job-tracker>
+        <name-node>${nameNode}</name-node>
+        <configuration>
+            <property>
+                <name>oozie.action.sharelib.for.spark</name>
+                <value>${oozieActionShareLibForSpark2}</value>
+            </property>
+        </configuration>
+    </global>
+
+    <start to="reset_outputpath"/>
+
+    <kill name="Kill">
+        <message>Action failed, error message[${wf:errorMessage(wf:lastErrorNode())}]</message>
+    </kill>
+
+    <action name="reset_outputpath">
+        <fs>
+            <delete path="${outputPath}"/>
+            <mkdir path="${outputPath}"/>
+        </fs>
+        <ok to="copy_entities"/>
+        <error to="Kill"/>
+    </action>
+
+    <fork name="copy_entities">
+        <path start="copy_relation"/>
+        <path start="copy_organization"/>
+        <path start="copy_projects"/>
+        <path start="copy_datasources"/>
+    </fork>
+
+    <action name="copy_relation">
+        <distcp xmlns="uri:oozie:distcp-action:0.2">
+            <arg>${nameNode}/${sourcePath}/relation</arg>
+            <arg>${nameNode}/${outputPath}/relation</arg>
+        </distcp>
+        <ok to="copy_wait"/>
+        <error to="Kill"/>
+    </action>
+
+    <action name="copy_organization">
+        <distcp xmlns="uri:oozie:distcp-action:0.2">
+            <arg>${nameNode}/${sourcePath}/organization</arg>
+            <arg>${nameNode}/${outputPath}/organization</arg>
+        </distcp>
+        <ok to="copy_wait"/>
+        <error to="Kill"/>
+    </action>
+
+    <action name="copy_projects">
+        <distcp xmlns="uri:oozie:distcp-action:0.2">
+            <arg>${nameNode}/${sourcePath}/project</arg>
+            <arg>${nameNode}/${outputPath}/project</arg>
+        </distcp>
+        <ok to="copy_wait"/>
+        <error to="Kill"/>
+    </action>
+
+    <action name="copy_datasources">
+        <distcp xmlns="uri:oozie:distcp-action:0.2">
+            <arg>${nameNode}/${sourcePath}/datasource</arg>
+            <arg>${nameNode}/${outputPath}/datasource</arg>
+        </distcp>
+        <ok to="copy_wait"/>
+        <error to="Kill"/>
+    </action>
+
+    <join name="copy_wait" to="prepare_result_communitylist"/>
+
+    <action name="prepare_result_communitylist">
+        <spark xmlns="uri:oozie:spark-action:0.2">
+            <master>yarn</master>
+            <mode>cluster</mode>
+            <name>Prepare-Community-Result-Organization</name>
+            <class>eu.dnetlib.dhp.resulttocommunityfromproject.PrepareResultCommunitySet</class>
+            <jar>dhp-enrichment-${projectVersion}.jar</jar>
+            <spark-opts>
+                --executor-cores=6
+                --executor-memory=5G
+                --conf spark.executor.memoryOverhead=3g
+                --conf spark.sql.shuffle.partitions=3284
+                --driver-memory=${sparkDriverMemory}
+                --conf spark.extraListeners=${spark2ExtraListeners}
+                --conf spark.sql.queryExecutionListeners=${spark2SqlQueryExecutionListeners}
+                --conf spark.yarn.historyServer.address=${spark2YarnHistoryServerAddress}
+                --conf spark.eventLog.dir=${nameNode}${spark2EventLogDir}
+
+                --conf spark.dynamicAllocation.maxExecutors=${spark2MaxExecutors}
+            </spark-opts>
+            <arg>--sourcePath</arg><arg>${sourcePath}/relation</arg>
+            <arg>--outputPath</arg><arg>${workingDir}/preparedInfo/resultCommunityList</arg>
+            <arg>--production</arg><arg>${production}</arg>
+        </spark>
+        <ok to="exec-propagation"/>
+        <error to="Kill"/>
+    </action>
+
+    <action name="exec-propagation">
+        <spark xmlns="uri:oozie:spark-action:0.2">
+            <master>yarn</master>
+            <mode>cluster</mode>
+            <name>community2resultfromproject</name>
+            <class>eu.dnetlib.dhp.resulttocommunityfromproject.SparkResultToCommunityFromProject</class>
+            <jar>dhp-enrichment-${projectVersion}.jar</jar>
+            <spark-opts>
+                --executor-cores=6
+                --executor-memory=5G
+                --conf spark.executor.memoryOverhead=3g
+                --conf spark.sql.shuffle.partitions=3284
+                --driver-memory=${sparkDriverMemory}
+                --conf spark.extraListeners=${spark2ExtraListeners}
+                --conf spark.sql.queryExecutionListeners=${spark2SqlQueryExecutionListeners}
+                --conf spark.yarn.historyServer.address=${spark2YarnHistoryServerAddress}
+                --conf spark.eventLog.dir=${nameNode}${spark2EventLogDir}
+                --conf spark.dynamicAllocation.maxExecutors=${spark2MaxExecutors}
+            </spark-opts>
+            <arg>--preparedInfoPath</arg><arg>${workingDir}/preparedInfo/resultCommunityList</arg>
+            <arg>--sourcePath</arg><arg>${sourcePath}/</arg>
+            <arg>--outputPath</arg><arg>${outputPath}/</arg>
+        </spark>
+        <ok to="End"/>
+        <error to="Kill"/>
+    </action>
+
+
+
+    <end name="End"/>
+
+</workflow-app>
--- a/dhp-workflows/dhp-enrichment/src/test/java/eu/dnetlib/dhp/api/QueryCommunityAPITest.java
+++ b/dhp-workflows/dhp-enrichment/src/test/java/eu/dnetlib/dhp/api/QueryCommunityAPITest.java
@ -0,0 +1,119 @@
+
+package eu.dnetlib.dhp.api;
+
+import java.util.List;
+
+import org.junit.jupiter.api.Assertions;
+import org.junit.jupiter.api.Test;
+
+import com.fasterxml.jackson.core.JsonProcessingException;
+import com.fasterxml.jackson.databind.ObjectMapper;
+
+import eu.dnetlib.dhp.api.QueryCommunityAPI;
+import eu.dnetlib.dhp.api.Utils;
+import eu.dnetlib.dhp.api.model.CommunityEntityMap;
+import eu.dnetlib.dhp.api.model.CommunityModel;
+import eu.dnetlib.dhp.api.model.CommunitySummary;
+import eu.dnetlib.dhp.api.model.DatasourceList;
+import eu.dnetlib.dhp.bulktag.community.Community;
+import eu.dnetlib.dhp.bulktag.community.CommunityConfiguration;
+
+/**
+ * @author miriam.baglioni
+ * @Date 06/10/23
+ */
+public class QueryCommunityAPITest {
+
+	@Test
+	void communityList() throws Exception {
+		String body = QueryCommunityAPI.communities(true);
+		new ObjectMapper()
+			.readValue(body, CommunitySummary.class)
+			.forEach(p -> {
+				try {
+					System.out.println(new ObjectMapper().writeValueAsString(p));
+				} catch (JsonProcessingException e) {
+					throw new RuntimeException(e);
+				}
+			});
+	}
+
+	@Test
+	void community() throws Exception {
+		String id = "dh-ch";
+		String body = QueryCommunityAPI.community(id, true);
+		System.out
+			.println(
+				new ObjectMapper()
+					.writeValueAsString(
+						new ObjectMapper()
+							.readValue(body, CommunityModel.class)));
+	}
+
+	@Test
+	void communityDatasource() throws Exception {
+		String id = "dh-ch";
+		String body = QueryCommunityAPI.communityDatasource(id, true);
+		new ObjectMapper()
+			.readValue(body, DatasourceList.class)
+			.forEach(ds -> {
+				try {
+					System.out.println(new ObjectMapper().writeValueAsString(ds));
+				} catch (JsonProcessingException e) {
+					throw new RuntimeException(e);
+				}
+			});
+		;
+	}
+
+	@Test
+	void validCommunities() throws Exception {
+		CommunityConfiguration cc = Utils.getCommunityConfiguration(true);
+		System.out.println(cc.getCommunities().keySet());
+		Community community = cc.getCommunities().get("aurora");
+		Assertions.assertEquals(0, community.getSubjects().size());
+		Assertions.assertEquals(null, community.getConstraints());
+		Assertions.assertEquals(null, community.getRemoveConstraints());
+		Assertions.assertEquals(2, community.getZenodoCommunities().size());
+		Assertions
+			.assertTrue(
+				community.getZenodoCommunities().stream().anyMatch(c -> c.equals("aurora-universities-network")));
+		Assertions
+			.assertTrue(community.getZenodoCommunities().stream().anyMatch(c -> c.equals("university-of-innsbruck")));
+		Assertions.assertEquals(35, community.getProviders().size());
+		Assertions
+			.assertEquals(
+				35, community.getProviders().stream().filter(p -> p.getSelectionConstraints() == null).count());
+
+	}
+
+	@Test
+	void eutopiaCommunityConfiguration() throws Exception {
+		CommunityConfiguration cc = Utils.getCommunityConfiguration(true);
+		System.out.println(cc.getCommunities().keySet());
+		Community community = cc.getCommunities().get("eutopia");
+		community.getProviders().forEach(p -> System.out.println(p.getOpenaireId()));
+	}
+
+	@Test
+	void getCommunityProjects() throws Exception {
+		CommunityEntityMap projectMap = Utils.getCommunityProjects(true);
+
+		Assertions
+			.assertTrue(
+				projectMap
+					.keySet()
+					.stream()
+					.allMatch(k -> k.startsWith("40|")));
+
+		System.out.println(projectMap);
+	}
+
+	@Test
+	void getCommunityOrganizations() throws Exception {
+		CommunityEntityMap organizationMap = Utils.getCommunityOrganization(true);
+		Assertions.assertTrue(organizationMap.keySet().stream().allMatch(k -> k.startsWith("20|")));
+
+	}
+
+}
--- a/dhp-workflows/dhp-enrichment/src/test/java/eu/dnetlib/dhp/bulktag/BulkTagJobTest.java
+++ b/dhp-workflows/dhp-enrichment/src/test/java/eu/dnetlib/dhp/bulktag/BulkTagJobTest.java
@ -6,6 +6,7 @@ import static eu.dnetlib.dhp.bulktag.community.TaggingConstants.ZENODO_COMMUNITY
 import java.io.IOException;
 import java.nio.file.Files;
 import java.nio.file.Path;
+import java.util.HashMap;
 import java.util.List;

 import org.apache.commons.io.FileUtils;
@ -98,14 +99,11 @@ public class BulkTagJobTest {
 		SparkBulkTagJob
 			.main(
 				new String[] {
-					"-isTest", Boolean.TRUE.toString(),
 					"-isSparkSessionManaged", Boolean.FALSE.toString(),
 					"-sourcePath",
-					getClass().getResource("/eu/dnetlib/dhp/bulktag/sample/dataset/no_updates").getPath(),
+					getClass().getResource("/eu/dnetlib/dhp/bulktag/sample/dataset/no_updates/").getPath(),
 					"-taggingConf", taggingConf,
-					"-resultTableName", "eu.dnetlib.dhp.schema.oaf.Dataset",
-					"-outputPath", workingDir.toString() + "/dataset",
-					"-isLookUpUrl", MOCK_IS_LOOK_UP_URL,
+					"-outputPath", workingDir.toString() + "/",
 					"-pathMap", pathMap
 				});

@ -133,19 +131,16 @@ public class BulkTagJobTest {
 	@Test
 	void bulktagBySubjectNoPreviousContextTest() throws Exception {
 		final String sourcePath = getClass()
-			.getResource("/eu/dnetlib/dhp/bulktag/sample/dataset/update_subject/nocontext")
+			.getResource("/eu/dnetlib/dhp/bulktag/sample/dataset/update_subject/nocontext/")
 			.getPath();
 		final String pathMap = BulkTagJobTest.pathMap;
 		SparkBulkTagJob
 			.main(
 				new String[] {
-					"-isTest", Boolean.TRUE.toString(),
 					"-isSparkSessionManaged", Boolean.FALSE.toString(),
 					"-sourcePath", sourcePath,
 					"-taggingConf", taggingConf,
-					"-resultTableName", "eu.dnetlib.dhp.schema.oaf.Dataset",
-					"-outputPath", workingDir.toString() + "/dataset",
-					"-isLookUpUrl", MOCK_IS_LOOK_UP_URL,
+					"-outputPath", workingDir.toString() + "/",
 					"-pathMap", pathMap
 				});

@ -230,19 +225,19 @@ public class BulkTagJobTest {
 	void bulktagBySubjectPreviousContextNoProvenanceTest() throws Exception {
 		final String sourcePath = getClass()
 			.getResource(
-				"/eu/dnetlib/dhp/bulktag/sample/dataset/update_subject/contextnoprovenance")
+				"/eu/dnetlib/dhp/bulktag/sample/dataset/update_subject/contextnoprovenance/")
 			.getPath();
 		final String pathMap = BulkTagJobTest.pathMap;
 		SparkBulkTagJob
 			.main(
 				new String[] {
-					"-isTest", Boolean.TRUE.toString(),
+
 					"-isSparkSessionManaged", Boolean.FALSE.toString(),
 					"-sourcePath", sourcePath,
 					"-taggingConf", taggingConf,
-					"-resultTableName", "eu.dnetlib.dhp.schema.oaf.Dataset",
-					"-outputPath", workingDir.toString() + "/dataset",
-					"-isLookUpUrl", MOCK_IS_LOOK_UP_URL,
+
+					"-outputPath", workingDir.toString() + "/",
+
 					"-pathMap", pathMap
 				});

@ -311,18 +306,18 @@ public class BulkTagJobTest {
 	@Test
 	void bulktagByDatasourceTest() throws Exception {
 		final String sourcePath = getClass()
-			.getResource("/eu/dnetlib/dhp/bulktag/sample/publication/update_datasource")
+			.getResource("/eu/dnetlib/dhp/bulktag/sample/publication/update_datasource/")
 			.getPath();
 		SparkBulkTagJob
 			.main(
 				new String[] {
-					"-isTest", Boolean.TRUE.toString(),
+
 					"-isSparkSessionManaged", Boolean.FALSE.toString(),
 					"-sourcePath", sourcePath,
 					"-taggingConf", taggingConf,
-					"-resultTableName", "eu.dnetlib.dhp.schema.oaf.Publication",
-					"-outputPath", workingDir.toString() + "/publication",
-					"-isLookUpUrl", MOCK_IS_LOOK_UP_URL,
+
+					"-outputPath", workingDir.toString() + "/",
+
 					"-pathMap", pathMap
 				});

@ -384,25 +379,25 @@ public class BulkTagJobTest {
 	void bulktagByZenodoCommunityTest() throws Exception {
 		final String sourcePath = getClass()
 			.getResource(
-				"/eu/dnetlib/dhp/bulktag/sample/otherresearchproduct/update_zenodocommunity")
+				"/eu/dnetlib/dhp/bulktag/sample/otherresearchproduct/update_zenodocommunity/")
 			.getPath();
 		SparkBulkTagJob
 			.main(
 				new String[] {
-					"-isTest", Boolean.TRUE.toString(),
+
 					"-isSparkSessionManaged", Boolean.FALSE.toString(),
 					"-sourcePath", sourcePath,
 					"-taggingConf", taggingConf,
-					"-resultTableName", "eu.dnetlib.dhp.schema.oaf.OtherResearchProduct",
-					"-outputPath", workingDir.toString() + "/orp",
-					"-isLookUpUrl", MOCK_IS_LOOK_UP_URL,
+
+					"-outputPath", workingDir.toString() + "/",
+
 					"-pathMap", pathMap
 				});

 		final JavaSparkContext sc = JavaSparkContext.fromSparkContext(spark.sparkContext());

 		JavaRDD<OtherResearchProduct> tmp = sc
-			.textFile(workingDir.toString() + "/orp")
+			.textFile(workingDir.toString() + "/otherresearchproduct")
 			.map(item -> OBJECT_MAPPER.readValue(item, OtherResearchProduct.class));

 		Assertions.assertEquals(10, tmp.count());
@ -505,18 +500,18 @@ public class BulkTagJobTest {
 	@Test
 	void bulktagBySubjectDatasourceTest() throws Exception {
 		final String sourcePath = getClass()
-			.getResource("/eu/dnetlib/dhp/bulktag/sample/dataset/update_subject_datasource")
+			.getResource("/eu/dnetlib/dhp/bulktag/sample/dataset/update_subject_datasource/")
 			.getPath();
 		SparkBulkTagJob
 			.main(
 				new String[] {
-					"-isTest", Boolean.TRUE.toString(),
+
 					"-isSparkSessionManaged", Boolean.FALSE.toString(),
 					"-sourcePath", sourcePath,
 					"-taggingConf", taggingConf,
-					"-resultTableName", "eu.dnetlib.dhp.schema.oaf.Dataset",
-					"-outputPath", workingDir.toString() + "/dataset",
-					"-isLookUpUrl", MOCK_IS_LOOK_UP_URL,
+
+					"-outputPath", workingDir.toString() + "/",
+
 					"-pathMap", pathMap
 				});

@ -636,14 +631,14 @@ public class BulkTagJobTest {
 		SparkBulkTagJob
 			.main(
 				new String[] {
-					"-isTest", Boolean.TRUE.toString(),
+
 					"-isSparkSessionManaged", Boolean.FALSE.toString(),
 					"-sourcePath",
-					getClass().getResource("/eu/dnetlib/dhp/bulktag/sample/software/software_10.json.gz").getPath(),
+					getClass().getResource("/eu/dnetlib/dhp/bulktag/sample/software/").getPath(),
 					"-taggingConf", taggingConf,
-					"-resultTableName", "eu.dnetlib.dhp.schema.oaf.Software",
-					"-outputPath", workingDir.toString() + "/software",
-					"-isLookUpUrl", MOCK_IS_LOOK_UP_URL,
+
+					"-outputPath", workingDir.toString() + "/",
+
 					"-pathMap", pathMap
 				});

@ -732,18 +727,18 @@ public class BulkTagJobTest {

 		final String sourcePath = getClass()
 			.getResource(
-				"/eu/dnetlib/dhp/bulktag/sample/dataset/update_datasourcewithconstraints")
+				"/eu/dnetlib/dhp/bulktag/sample/dataset/update_datasourcewithconstraints/")
 			.getPath();
 		SparkBulkTagJob
 			.main(
 				new String[] {
-					"-isTest", Boolean.TRUE.toString(),
+
 					"-isSparkSessionManaged", Boolean.FALSE.toString(),
 					"-sourcePath", sourcePath,
 					"-taggingConf", taggingConf,
-					"-resultTableName", "eu.dnetlib.dhp.schema.oaf.Dataset",
-					"-outputPath", workingDir.toString() + "/dataset",
-					"-isLookUpUrl", MOCK_IS_LOOK_UP_URL,
+
+					"-outputPath", workingDir.toString() + "/",
+
 					"-pathMap", pathMap
 				});

@ -774,19 +769,19 @@ public class BulkTagJobTest {
 	void bulkTagOtherJupyter() throws Exception {
 		final String sourcePath = getClass()
 			.getResource(
-				"/eu/dnetlib/dhp/eosctag/jupyter/otherresearchproduct")
+				"/eu/dnetlib/dhp/eosctag/jupyter/")
 			.getPath();

 		SparkBulkTagJob
 			.main(
 				new String[] {
-					"-isTest", Boolean.TRUE.toString(),
+
 					"-isSparkSessionManaged", Boolean.FALSE.toString(),
 					"-sourcePath", sourcePath,
 					"-taggingConf", taggingConf,
-					"-resultTableName", "eu.dnetlib.dhp.schema.oaf.OtherResearchProduct",
-					"-outputPath", workingDir.toString() + "/otherresearchproduct",
-					"-isLookUpUrl", MOCK_IS_LOOK_UP_URL,
+
+					"-outputPath", workingDir.toString() + "/",
+
 					"-pathMap", pathMap
 				});

@ -829,18 +824,18 @@ public class BulkTagJobTest {
 	public void bulkTagDatasetJupyter() throws Exception {
 		final String sourcePath = getClass()
 			.getResource(
-				"/eu/dnetlib/dhp/eosctag/jupyter/dataset")
+				"/eu/dnetlib/dhp/eosctag/jupyter/")
 			.getPath();
 		SparkBulkTagJob
 			.main(
 				new String[] {
-					"-isTest", Boolean.TRUE.toString(),
+
 					"-isSparkSessionManaged", Boolean.FALSE.toString(),
 					"-sourcePath", sourcePath,
 					"-taggingConf", taggingConf,
-					"-resultTableName", "eu.dnetlib.dhp.schema.oaf.Dataset",
-					"-outputPath", workingDir.toString() + "/dataset",
-					"-isLookUpUrl", MOCK_IS_LOOK_UP_URL,
+
+					"-outputPath", workingDir.toString() + "/",
+
 					"-pathMap", pathMap
 				});

@ -878,18 +873,18 @@ public class BulkTagJobTest {

 		final String sourcePath = getClass()
 			.getResource(
-				"/eu/dnetlib/dhp/eosctag/jupyter/software")
+				"/eu/dnetlib/dhp/eosctag/jupyter/")
 			.getPath();
 		SparkBulkTagJob
 			.main(
 				new String[] {
-					"-isTest", Boolean.TRUE.toString(),
+
 					"-isSparkSessionManaged", Boolean.FALSE.toString(),
 					"-sourcePath", sourcePath,
 					"-taggingConf", taggingConf,
-					"-resultTableName", "eu.dnetlib.dhp.schema.oaf.Software",
-					"-outputPath", workingDir.toString() + "/software",
-					"-isLookUpUrl", MOCK_IS_LOOK_UP_URL,
+
+					"-outputPath", workingDir.toString() + "/",
+
 					"-pathMap", pathMap
 				});

@ -1096,18 +1091,18 @@ public class BulkTagJobTest {
 	void galaxyOtherTest() throws Exception {
 		final String sourcePath = getClass()
 			.getResource(
-				"/eu/dnetlib/dhp/eosctag/galaxy/otherresearchproduct")
+				"/eu/dnetlib/dhp/eosctag/galaxy/")
 			.getPath();
 		SparkBulkTagJob
 			.main(
 				new String[] {
-					"-isTest", Boolean.TRUE.toString(),
+
 					"-isSparkSessionManaged", Boolean.FALSE.toString(),
 					"-sourcePath", sourcePath,
 					"-taggingConf", taggingConf,
-					"-resultTableName", "eu.dnetlib.dhp.schema.oaf.OtherResearchProduct",
-					"-outputPath", workingDir.toString() + "/otherresearchproduct",
-					"-isLookUpUrl", MOCK_IS_LOOK_UP_URL,
+
+					"-outputPath", workingDir.toString() + "/",
+
 					"-pathMap", pathMap
 				});

@ -1214,18 +1209,18 @@ public class BulkTagJobTest {
 	void galaxySoftwareTest() throws Exception {
 		final String sourcePath = getClass()
 			.getResource(
-				"/eu/dnetlib/dhp/eosctag/galaxy/software")
+				"/eu/dnetlib/dhp/eosctag/galaxy/")
 			.getPath();
 		SparkBulkTagJob
 			.main(
 				new String[] {
-					"-isTest", Boolean.TRUE.toString(),
+
 					"-isSparkSessionManaged", Boolean.FALSE.toString(),
 					"-sourcePath", sourcePath,
 					"-taggingConf", taggingConf,
-					"-resultTableName", "eu.dnetlib.dhp.schema.oaf.Software",
-					"-outputPath", workingDir.toString() + "/software",
-					"-isLookUpUrl", MOCK_IS_LOOK_UP_URL,
+
+					"-outputPath", workingDir.toString() + "/",
+
 					"-pathMap", pathMap
 				});

@ -1333,19 +1328,19 @@ public class BulkTagJobTest {
 	void twitterDatasetTest() throws Exception {
 		final String sourcePath = getClass()
 			.getResource(
-				"/eu/dnetlib/dhp/eosctag/twitter/dataset")
+				"/eu/dnetlib/dhp/eosctag/twitter/")
 			.getPath();

 		SparkBulkTagJob
 			.main(
 				new String[] {
-					"-isTest", Boolean.TRUE.toString(),
+
 					"-isSparkSessionManaged", Boolean.FALSE.toString(),
 					"-sourcePath", sourcePath,
 					"-taggingConf", taggingConf,
-					"-resultTableName", "eu.dnetlib.dhp.schema.oaf.Dataset",
-					"-outputPath", workingDir.toString() + "/dataset",
-					"-isLookUpUrl", MOCK_IS_LOOK_UP_URL,
+
+					"-outputPath", workingDir.toString() + "/",
+
 					"-pathMap", pathMap
 				});

@ -1373,19 +1368,19 @@ public class BulkTagJobTest {
 	void twitterOtherTest() throws Exception {
 		final String sourcePath = getClass()
 			.getResource(
-				"/eu/dnetlib/dhp/eosctag/twitter/otherresearchproduct")
+				"/eu/dnetlib/dhp/eosctag/twitter/")
 			.getPath();

 		SparkBulkTagJob
 			.main(
 				new String[] {
-					"-isTest", Boolean.TRUE.toString(),
+
 					"-isSparkSessionManaged", Boolean.FALSE.toString(),
 					"-sourcePath", sourcePath,
 					"-taggingConf", taggingConf,
-					"-resultTableName", "eu.dnetlib.dhp.schema.oaf.OtherResearchProduct",
-					"-outputPath", workingDir.toString() + "/otherresearchproduct",
-					"-isLookUpUrl", MOCK_IS_LOOK_UP_URL,
+
+					"-outputPath", workingDir.toString() + "/",
+
 					"-pathMap", pathMap
 				});

@ -1418,19 +1413,19 @@ public class BulkTagJobTest {
 	void twitterSoftwareTest() throws Exception {
 		final String sourcePath = getClass()
 			.getResource(
-				"/eu/dnetlib/dhp/eosctag/twitter/software")
+				"/eu/dnetlib/dhp/eosctag/twitter/")
 			.getPath();

 		SparkBulkTagJob
 			.main(
 				new String[] {
-					"-isTest", Boolean.TRUE.toString(),
+
 					"-isSparkSessionManaged", Boolean.FALSE.toString(),
 					"-sourcePath", sourcePath,
 					"-taggingConf", taggingConf,
-					"-resultTableName", "eu.dnetlib.dhp.schema.oaf.Software",
-					"-outputPath", workingDir.toString() + "/software",
-					"-isLookUpUrl", MOCK_IS_LOOK_UP_URL,
+
+					"-outputPath", workingDir.toString() + "/",
+
 					"-pathMap", pathMap
 				});

@ -1455,19 +1450,19 @@ public class BulkTagJobTest {
 	void EoscContextTagTest() throws Exception {
 		final String sourcePath = getClass()
 			.getResource(
-				"/eu/dnetlib/dhp/bulktag/eosc/dataset/dataset_10.json")
+				"/eu/dnetlib/dhp/bulktag/eosc/dataset/")
 			.getPath();

 		SparkBulkTagJob
 			.main(
 				new String[] {
-					"-isTest", Boolean.TRUE.toString(),
+
 					"-isSparkSessionManaged", Boolean.FALSE.toString(),
 					"-sourcePath", sourcePath,
 					"-taggingConf", taggingConf,
-					"-resultTableName", "eu.dnetlib.dhp.schema.oaf.Dataset",
-					"-outputPath", workingDir.toString() + "/dataset",
-					"-isLookUpUrl", MOCK_IS_LOOK_UP_URL,
+
+					"-outputPath", workingDir.toString() + "/",
+
 					"-pathMap", pathMap
 				});

@ -1533,16 +1528,16 @@ public class BulkTagJobTest {
 		SparkBulkTagJob
 			.main(
 				new String[] {
-					"-isTest", Boolean.TRUE.toString(),
+
 					"-isSparkSessionManaged", Boolean.FALSE.toString(),
 					"-sourcePath",
 					getClass()
-						.getResource("/eu/dnetlib/dhp/bulktag/sample/dataset/update_datasourcewithconstraints")
+						.getResource("/eu/dnetlib/dhp/bulktag/sample/dataset/update_datasourcewithconstraints/")
 						.getPath(),
 					"-taggingConf", taggingConf,
-					"-resultTableName", "eu.dnetlib.dhp.schema.oaf.Dataset",
-					"-outputPath", workingDir.toString() + "/dataset",
-					"-isLookUpUrl", MOCK_IS_LOOK_UP_URL,
+
+					"-outputPath", workingDir.toString() + "/",
+
 					"-pathMap", pathMap
 				});
 		final JavaSparkContext sc = JavaSparkContext.fromSparkContext(spark.sparkContext());
@ -1568,4 +1563,42 @@ public class BulkTagJobTest {

 	}

+	@Test
+	void newConfTest() throws Exception {
+		final String pathMap = BulkTagJobTest.pathMap;
+		SparkBulkTagJob
+			.main(
+				new String[] {
+
+					"-isSparkSessionManaged", Boolean.FALSE.toString(),
+					"-sourcePath",
+					getClass().getResource("/eu/dnetlib/dhp/bulktag/sample/dataset/no_updates/").getPath(),
+					"-taggingConf", taggingConf,
+
+					"-outputPath", workingDir.toString() + "/",
+					"-production", Boolean.TRUE.toString(),
+					"-pathMap", pathMap
+				});
+
+		final JavaSparkContext sc = JavaSparkContext.fromSparkContext(spark.sparkContext());
+
+		JavaRDD<Dataset> tmp = sc
+			.textFile(workingDir.toString() + "/dataset")
+			.map(item -> OBJECT_MAPPER.readValue(item, Dataset.class));
+
+		Assertions.assertEquals(10, tmp.count());
+		org.apache.spark.sql.Dataset<Dataset> verificationDataset = spark
+			.createDataset(tmp.rdd(), Encoders.bean(Dataset.class));
+
+		verificationDataset.createOrReplaceTempView("dataset");
+
+		String query = "select id, MyT.id community "
+			+ "from dataset "
+			+ "lateral view explode(context) c as MyT "
+			+ "lateral view explode(MyT.datainfo) d as MyD "
+			+ "where MyD.inferenceprovenance = 'bulktagging'";
+
+		Assertions.assertEquals(0, spark.sql(query).count());
+	}
+
 }
--- a/dhp-workflows/dhp-enrichment/src/test/java/eu/dnetlib/dhp/bulktag/CommunityConfigurationFactoryTest.java
+++ b/dhp-workflows/dhp-enrichment/src/test/java/eu/dnetlib/dhp/bulktag/CommunityConfigurationFactoryTest.java
@ -47,7 +47,7 @@ class CommunityConfigurationFactoryTest {
 		sc.setVerb("not_contains");
 		sc.setField("contributor");
 		sc.setValue("DARIAH");
-		sc.setSelection(resolver.getSelectionCriteria(sc.getVerb(), sc.getValue()));
+		sc.setSelection(resolver);// .getSelectionCriteria(sc.getVerb(), sc.getValue()));
 		String metadata = "This work has been partially supported by DARIAH-EU infrastructure";
 		Assertions.assertFalse(sc.verifyCriteria(metadata));
 	}
--- a/dhp-workflows/dhp-enrichment/src/test/java/eu/dnetlib/dhp/resulttocommunityfromorganization/PrepareAssocTest.java
+++ b/dhp-workflows/dhp-enrichment/src/test/java/eu/dnetlib/dhp/resulttocommunityfromorganization/PrepareAssocTest.java
@ -0,0 +1,95 @@
+
+package eu.dnetlib.dhp.resulttocommunityfromorganization;
+
+import java.io.IOException;
+import java.nio.file.Files;
+import java.nio.file.Path;
+
+import org.apache.commons.io.FileUtils;
+import org.apache.commons.io.IOUtils;
+import org.apache.spark.SparkConf;
+import org.apache.spark.api.java.JavaRDD;
+import org.apache.spark.api.java.JavaSparkContext;
+import org.apache.spark.sql.Encoders;
+import org.apache.spark.sql.SparkSession;
+import org.junit.jupiter.api.AfterAll;
+import org.junit.jupiter.api.Assertions;
+import org.junit.jupiter.api.BeforeAll;
+import org.junit.jupiter.api.Test;
+import org.slf4j.Logger;
+import org.slf4j.LoggerFactory;
+
+import com.fasterxml.jackson.databind.ObjectMapper;
+import com.google.gson.Gson;
+
+import eu.dnetlib.dhp.api.Utils;
+import eu.dnetlib.dhp.api.model.CommunityEntityMap;
+import eu.dnetlib.dhp.bulktag.BulkTagJobTest;
+import eu.dnetlib.dhp.bulktag.SparkBulkTagJob;
+import eu.dnetlib.dhp.schema.oaf.Dataset;
+
+/**
+ * @author miriam.baglioni
+ * @Date 13/10/23
+ */
+public class PrepareAssocTest {
+
+	private static SparkSession spark;
+
+	private static Path workingDir;
+
+	private static final Logger log = LoggerFactory.getLogger(PrepareAssocTest.class);
+
+	@BeforeAll
+	public static void beforeAll() throws IOException {
+		workingDir = Files.createTempDirectory(BulkTagJobTest.class.getSimpleName());
+		log.info("using work dir {}", workingDir);
+
+		SparkConf conf = new SparkConf();
+		conf.setAppName(BulkTagJobTest.class.getSimpleName());
+
+		conf.setMaster("local[*]");
+		conf.set("spark.driver.host", "localhost");
+		conf.set("hive.metastore.local", "true");
+		conf.set("spark.ui.enabled", "false");
+		conf.set("spark.sql.warehouse.dir", workingDir.toString());
+		conf.set("hive.metastore.warehouse.dir", workingDir.resolve("warehouse").toString());
+
+		spark = SparkSession
+			.builder()
+			.appName(PrepareAssocTest.class.getSimpleName())
+			.config(conf)
+			.getOrCreate();
+	}
+
+	@AfterAll
+	public static void afterAll() throws IOException {
+		FileUtils.deleteDirectory(workingDir.toFile());
+		spark.stop();
+	}
+
+	@Test
+	void test1() throws Exception {
+
+		PrepareResultCommunitySet
+			.main(
+				new String[] {
+
+					"-isSparkSessionManaged", Boolean.FALSE.toString(),
+					"-sourcePath",
+					getClass().getResource("/eu/dnetlib/dhp/resulttocommunityfromorganization/relation/").getPath(),
+					"-outputPath", workingDir.toString() + "/prepared",
+					"-production", Boolean.TRUE.toString(),
+					"-hive_metastore_uris", ""
+				});
+
+		final JavaSparkContext sc = JavaSparkContext.fromSparkContext(spark.sparkContext());
+
+		JavaRDD<ResultCommunityList> tmp = sc
+			.textFile(workingDir.toString() + "/prepared")
+			.map(item -> new ObjectMapper().readValue(item, ResultCommunityList.class));
+
+		tmp.foreach(r -> System.out.println(new ObjectMapper().writeValueAsString(r)));
+	}
+
+}
--- a/dhp-workflows/dhp-enrichment/src/test/java/eu/dnetlib/dhp/resulttocommunityfromorganization/ResultToCommunityJobTest.java
+++ b/dhp-workflows/dhp-enrichment/src/test/java/eu/dnetlib/dhp/resulttocommunityfromorganization/ResultToCommunityJobTest.java
@ -78,7 +78,7 @@ public class ResultToCommunityJobTest {
 						.getResource("/eu/dnetlib/dhp/resulttocommunityfromorganization/sample")
 						.getPath(),
 					"-hive_metastore_uris", "",
-					"-saveGraph", "true",
+
 					"-resultTableName", "eu.dnetlib.dhp.schema.oaf.Dataset",
 					"-outputPath", workingDir.toString() + "/dataset",
 					"-preparedInfoPath", preparedInfoPath
--- a/dhp-workflows/dhp-enrichment/src/test/java/eu/dnetlib/dhp/resulttocommunityfromproject/PrepareAssocTest.java
+++ b/dhp-workflows/dhp-enrichment/src/test/java/eu/dnetlib/dhp/resulttocommunityfromproject/PrepareAssocTest.java
@ -0,0 +1,88 @@
+
+package eu.dnetlib.dhp.resulttocommunityfromproject;
+
+import java.io.IOException;
+import java.nio.file.Files;
+import java.nio.file.Path;
+
+import org.apache.commons.io.FileUtils;
+import org.apache.spark.SparkConf;
+import org.apache.spark.api.java.JavaRDD;
+import org.apache.spark.api.java.JavaSparkContext;
+import org.apache.spark.sql.SparkSession;
+import org.junit.jupiter.api.AfterAll;
+import org.junit.jupiter.api.BeforeAll;
+import org.junit.jupiter.api.Test;
+import org.slf4j.Logger;
+import org.slf4j.LoggerFactory;
+
+import com.fasterxml.jackson.databind.ObjectMapper;
+
+import eu.dnetlib.dhp.bulktag.BulkTagJobTest;
+import eu.dnetlib.dhp.resulttocommunityfromorganization.ResultCommunityList;
+
+/**
+ * @author miriam.baglioni
+ * @Date 13/10/23
+ */
+public class PrepareAssocTest {
+
+	private static SparkSession spark;
+
+	private static Path workingDir;
+
+	private static final Logger log = LoggerFactory.getLogger(PrepareAssocTest.class);
+
+	@BeforeAll
+	public static void beforeAll() throws IOException {
+		workingDir = Files.createTempDirectory(BulkTagJobTest.class.getSimpleName());
+		log.info("using work dir {}", workingDir);
+
+		SparkConf conf = new SparkConf();
+		conf.setAppName(BulkTagJobTest.class.getSimpleName());
+
+		conf.setMaster("local[*]");
+		conf.set("spark.driver.host", "localhost");
+		conf.set("hive.metastore.local", "true");
+		conf.set("spark.ui.enabled", "false");
+		conf.set("spark.sql.warehouse.dir", workingDir.toString());
+		conf.set("hive.metastore.warehouse.dir", workingDir.resolve("warehouse").toString());
+
+		spark = SparkSession
+			.builder()
+			.appName(PrepareAssocTest.class.getSimpleName())
+			.config(conf)
+			.getOrCreate();
+	}
+
+	@AfterAll
+	public static void afterAll() throws IOException {
+		FileUtils.deleteDirectory(workingDir.toFile());
+		spark.stop();
+	}
+
+	@Test
+	void test1() throws Exception {
+
+		PrepareResultCommunitySet
+			.main(
+				new String[] {
+
+					"-isSparkSessionManaged", Boolean.FALSE.toString(),
+					"-sourcePath",
+					getClass().getResource("/eu/dnetlib/dhp/resulttocommunityfromproject/relation/").getPath(),
+					"-outputPath", workingDir.toString() + "/prepared",
+					"-production", Boolean.TRUE.toString(),
+					"-hive_metastore_uris", ""
+				});
+
+		final JavaSparkContext sc = JavaSparkContext.fromSparkContext(spark.sparkContext());
+
+		JavaRDD<ResultProjectList> tmp = sc
+			.textFile(workingDir.toString() + "/prepared")
+			.map(item -> new ObjectMapper().readValue(item, ResultProjectList.class));
+
+		tmp.foreach(r -> System.out.println(new ObjectMapper().writeValueAsString(r)));
+	}
+
+}
--- a/dhp-workflows/dhp-enrichment/src/test/java/eu/dnetlib/dhp/resulttocommunityfromproject/ResultToCommunityJobTest.java
+++ b/dhp-workflows/dhp-enrichment/src/test/java/eu/dnetlib/dhp/resulttocommunityfromproject/ResultToCommunityJobTest.java
@ -0,0 +1,323 @@
+
+package eu.dnetlib.dhp.resulttocommunityfromproject;
+
+import static org.apache.spark.sql.functions.desc;
+
+import java.io.IOException;
+import java.nio.file.Files;
+import java.nio.file.Path;
+
+import org.apache.commons.io.FileUtils;
+import org.apache.spark.SparkConf;
+import org.apache.spark.api.java.JavaRDD;
+import org.apache.spark.api.java.JavaSparkContext;
+import org.apache.spark.sql.Encoders;
+import org.apache.spark.sql.Row;
+import org.apache.spark.sql.SparkSession;
+import org.junit.jupiter.api.AfterAll;
+import org.junit.jupiter.api.Assertions;
+import org.junit.jupiter.api.BeforeAll;
+import org.junit.jupiter.api.Test;
+import org.slf4j.Logger;
+import org.slf4j.LoggerFactory;
+
+import com.fasterxml.jackson.databind.ObjectMapper;
+
+import eu.dnetlib.dhp.orcidtoresultfromsemrel.OrcidPropagationJobTest;
+import eu.dnetlib.dhp.resulttocommunityfromorganization.SparkResultToCommunityFromOrganizationJob;
+import eu.dnetlib.dhp.schema.oaf.Dataset;
+
+public class ResultToCommunityJobTest {
+
+	private static final Logger log = LoggerFactory.getLogger(ResultToCommunityJobTest.class);
+
+	private static final ObjectMapper OBJECT_MAPPER = new ObjectMapper();
+
+	private static SparkSession spark;
+
+	private static Path workingDir;
+
+	@BeforeAll
+	public static void beforeAll() throws IOException {
+		workingDir = Files.createTempDirectory(ResultToCommunityJobTest.class.getSimpleName());
+		log.info("using work dir {}", workingDir);
+
+		SparkConf conf = new SparkConf();
+		conf.setAppName(ResultToCommunityJobTest.class.getSimpleName());
+
+		conf.setMaster("local[*]");
+		conf.set("spark.driver.host", "localhost");
+		conf.set("hive.metastore.local", "true");
+		conf.set("spark.ui.enabled", "false");
+		conf.set("spark.sql.warehouse.dir", workingDir.toString());
+		conf.set("hive.metastore.warehouse.dir", workingDir.resolve("warehouse").toString());
+
+		spark = SparkSession
+			.builder()
+			.appName(OrcidPropagationJobTest.class.getSimpleName())
+			.config(conf)
+			.getOrCreate();
+	}
+
+	@AfterAll
+	public static void afterAll() throws IOException {
+		FileUtils.deleteDirectory(workingDir.toFile());
+		spark.stop();
+	}
+
+	@Test
+	void testSparkResultToCommunityFromProjectJob() throws Exception {
+		final String preparedInfoPath = getClass()
+			.getResource("/eu/dnetlib/dhp/resulttocommunityfromproject/preparedInfo")
+			.getPath();
+		SparkResultToCommunityFromProject
+			.main(
+				new String[] {
+
+					"-isSparkSessionManaged", Boolean.FALSE.toString(),
+					"-sourcePath", getClass()
+						.getResource("/eu/dnetlib/dhp/resulttocommunityfromproject/sample/")
+						.getPath(),
+
+					"-outputPath", workingDir.toString() + "/",
+					"-preparedInfoPath", preparedInfoPath
+				});
+
+		final JavaSparkContext sc = JavaSparkContext.fromSparkContext(spark.sparkContext());
+
+		JavaRDD<Dataset> tmp = sc
+			.textFile(workingDir.toString() + "/dataset")
+			.map(item -> OBJECT_MAPPER.readValue(item, Dataset.class));
+
+		tmp.foreach(d -> System.out.println(new ObjectMapper().writeValueAsString(d)));
+//		Assertions.assertEquals(10, tmp.count());
+//		org.apache.spark.sql.Dataset<Dataset> verificationDataset = spark
+//			.createDataset(tmp.rdd(), Encoders.bean(Dataset.class));
+//
+//		verificationDataset.createOrReplaceTempView("dataset");
+//
+//		String query = "select id, MyT.id community "
+//			+ "from dataset "
+//			+ "lateral view explode(context) c as MyT "
+//			+ "lateral view explode(MyT.datainfo) d as MyD "
+//			+ "where MyD.inferenceprovenance = 'propagation'";
+//
+//		org.apache.spark.sql.Dataset<Row> resultExplodedProvenance = spark.sql(query);
+//		Assertions.assertEquals(5, resultExplodedProvenance.count());
+//		Assertions
+//			.assertEquals(
+//				0,
+//				resultExplodedProvenance
+//					.filter("id = '50|dedup_wf_001::afaf128022d29872c4dad402b2db04fe'")
+//					.count());
+//		Assertions
+//			.assertEquals(
+//				1,
+//				resultExplodedProvenance
+//					.filter("id = '50|dedup_wf_001::3f62cfc27024d564ea86760c494ba93b'")
+//					.count());
+//		Assertions
+//			.assertEquals(
+//				"beopen",
+//				resultExplodedProvenance
+//					.select("community")
+//					.where(
+//						resultExplodedProvenance
+//							.col("id")
+//							.equalTo(
+//								"50|dedup_wf_001::3f62cfc27024d564ea86760c494ba93b"))
+//					.collectAsList()
+//					.get(0)
+//					.getString(0));
+//
+//		Assertions
+//			.assertEquals(
+//				2,
+//				resultExplodedProvenance
+//					.filter("id = '50|od________18::8887b1df8b563c4ea851eb9c882c9d7b'")
+//					.count());
+//		Assertions
+//			.assertEquals(
+//				"mes",
+//				resultExplodedProvenance
+//					.select("community")
+//					.where(
+//						resultExplodedProvenance
+//							.col("id")
+//							.equalTo(
+//								"50|od________18::8887b1df8b563c4ea851eb9c882c9d7b"))
+//					.sort(desc("community"))
+//					.collectAsList()
+//					.get(0)
+//					.getString(0));
+//		Assertions
+//			.assertEquals(
+//				"euromarine",
+//				resultExplodedProvenance
+//					.select("community")
+//					.where(
+//						resultExplodedProvenance
+//							.col("id")
+//							.equalTo(
+//								"50|od________18::8887b1df8b563c4ea851eb9c882c9d7b"))
+//					.sort(desc("community"))
+//					.collectAsList()
+//					.get(1)
+//					.getString(0));
+//
+//		Assertions
+//			.assertEquals(
+//				1,
+//				resultExplodedProvenance
+//					.filter("id = '50|doajarticles::8d817039a63710fcf97e30f14662c6c8'")
+//					.count());
+//		Assertions
+//			.assertEquals(
+//				"mes",
+//				resultExplodedProvenance
+//					.select("community")
+//					.where(
+//						resultExplodedProvenance
+//							.col("id")
+//							.equalTo(
+//								"50|doajarticles::8d817039a63710fcf97e30f14662c6c8"))
+//					.sort(desc("community"))
+//					.collectAsList()
+//					.get(0)
+//					.getString(0));
+//
+//		Assertions
+//			.assertEquals(
+//				1,
+//				resultExplodedProvenance
+//					.filter("id = '50|doajarticles::3c98f0632f1875b4979e552ba3aa01e6'")
+//					.count());
+//		Assertions
+//			.assertEquals(
+//				"mes",
+//				resultExplodedProvenance
+//					.select("community")
+//					.where(
+//						resultExplodedProvenance
+//							.col("id")
+//							.equalTo(
+//								"50|doajarticles::3c98f0632f1875b4979e552ba3aa01e6"))
+//					.sort(desc("community"))
+//					.collectAsList()
+//					.get(0)
+//					.getString(0));
+//
+//		query = "select id, MyT.id community "
+//			+ "from dataset "
+//			+ "lateral view explode(context) c as MyT "
+//			+ "lateral view explode(MyT.datainfo) d as MyD ";
+//
+//		org.apache.spark.sql.Dataset<Row> resultCommunityId = spark.sql(query);
+//
+//		Assertions.assertEquals(10, resultCommunityId.count());
+//
+//		Assertions
+//			.assertEquals(
+//				1,
+//				resultCommunityId
+//					.filter("id = '50|dedup_wf_001::afaf128022d29872c4dad402b2db04fe'")
+//					.count());
+//		Assertions
+//			.assertEquals(
+//				"beopen",
+//				resultCommunityId
+//					.select("community")
+//					.where(
+//						resultCommunityId
+//							.col("id")
+//							.equalTo(
+//								"50|dedup_wf_001::afaf128022d29872c4dad402b2db04fe"))
+//					.collectAsList()
+//					.get(0)
+//					.getString(0));
+//
+//		Assertions
+//			.assertEquals(
+//				1,
+//				resultCommunityId
+//					.filter("id = '50|dedup_wf_001::3f62cfc27024d564ea86760c494ba93b'")
+//					.count());
+//
+//		Assertions
+//			.assertEquals(
+//				3,
+//				resultCommunityId
+//					.filter("id = '50|od________18::8887b1df8b563c4ea851eb9c882c9d7b'")
+//					.count());
+//		Assertions
+//			.assertEquals(
+//				"beopen",
+//				resultCommunityId
+//					.select("community")
+//					.where(
+//						resultCommunityId
+//							.col("id")
+//							.equalTo(
+//								"50|od________18::8887b1df8b563c4ea851eb9c882c9d7b"))
+//					.sort(desc("community"))
+//					.collectAsList()
+//					.get(2)
+//					.getString(0));
+//
+//		Assertions
+//			.assertEquals(
+//				2,
+//				resultCommunityId
+//					.filter("id = '50|doajarticles::8d817039a63710fcf97e30f14662c6c8'")
+//					.count());
+//		Assertions
+//			.assertEquals(
+//				"euromarine",
+//				resultCommunityId
+//					.select("community")
+//					.where(
+//						resultCommunityId
+//							.col("id")
+//							.equalTo(
+//								"50|doajarticles::8d817039a63710fcf97e30f14662c6c8"))
+//					.sort(desc("community"))
+//					.collectAsList()
+//					.get(1)
+//					.getString(0));
+//
+//		Assertions
+//			.assertEquals(
+//				3,
+//				resultCommunityId
+//					.filter("id = '50|doajarticles::3c98f0632f1875b4979e552ba3aa01e6'")
+//					.count());
+//		Assertions
+//			.assertEquals(
+//				"euromarine",
+//				resultCommunityId
+//					.select("community")
+//					.where(
+//						resultCommunityId
+//							.col("id")
+//							.equalTo(
+//								"50|doajarticles::3c98f0632f1875b4979e552ba3aa01e6"))
+//					.sort(desc("community"))
+//					.collectAsList()
+//					.get(2)
+//					.getString(0));
+//		Assertions
+//			.assertEquals(
+//				"ni",
+//				resultCommunityId
+//					.select("community")
+//					.where(
+//						resultCommunityId
+//							.col("id")
+//							.equalTo(
+//								"50|doajarticles::3c98f0632f1875b4979e552ba3aa01e6"))
+//					.sort(desc("community"))
+//					.collectAsList()
+//					.get(0)
+//					.getString(0));
+	}
+}
--- a/dhp-workflows/dhp-enrichment/src/test/resources/eu/dnetlib/dhp/bulktag/eosc/dataset/dataset_10.json
+++ b/dhp-workflows/dhp-enrichment/src/test/resources/eu/dnetlib/dhp/bulktag/eosc/dataset/dataset_10.json
--- a/dhp-workflows/dhp-enrichment/src/test/resources/eu/dnetlib/dhp/bulktag/eosc/dataset/otherresearchproduct
+++ b/dhp-workflows/dhp-enrichment/src/test/resources/eu/dnetlib/dhp/bulktag/eosc/dataset/otherresearchproduct
--- a/dhp-workflows/dhp-enrichment/src/test/resources/eu/dnetlib/dhp/bulktag/eosc/dataset/publication
+++ b/dhp-workflows/dhp-enrichment/src/test/resources/eu/dnetlib/dhp/bulktag/eosc/dataset/publication
--- a/Show More
+++ b/Show More
Author	SHA1	Message	Date
Miriam Baglioni	a4214ced1e	fixing issue on propagation organization. added --config to workflow definition. added oozie_app to communtiy project	2023-10-20 10:14:20 +02:00
Miriam Baglioni	f1b898c6b4	mergin with branch beta	2023-10-19 09:04:35 +02:00
Claudio Atzori	3b1c8b9fbd	Merge pull request 'FIX: GroupEntitiesSparkJob deletes whole graph outputPath instead of its temporary folder' (#351 ) from fix_consistency_missing_rels into beta Reviewed-on: D-Net/dnet-hadoop#351	2023-10-17 08:40:23 +02:00
Claudio Atzori	1d594eaffd	Merge branch 'beta' into fix_consistency_missing_rels	2023-10-17 08:40:07 +02:00
Giambattista Bloisi	0e44b037a5	FIX: GroupEntitiesSparkJob deletes whole graph outputPath instead of its temporary folder	2023-10-17 07:54:01 +02:00
Claudio Atzori	389e3fcc59	Merge pull request '[dedup] use common `saveParquet` and `save` methods to ensure outputs are compressed' (#349 ) from fix_dedup_not_compressed into beta Reviewed-on: D-Net/dnet-hadoop#349	2023-10-16 11:56:18 +02:00
Sandro La Bruzzo	a5a89a702f	new spark parrameter updated	2023-10-16 11:46:12 +02:00
Miriam Baglioni	159388f9c2	testing and fix some issues	2023-10-16 11:26:07 +02:00
Claudio Atzori	03670bb9ce	[dedup] use common saveParquet and save methods to ensure outputs are compressed	2023-10-16 10:55:47 +02:00
Claudio Atzori	6cf64d5d8b	[SWH] renamed 'Software Heritage Identifier' to 'Software Hash Identifier'	2023-10-13 10:09:26 +02:00
Claudio Atzori	76447958bb	cleanup & docs	2023-10-12 12:23:20 +02:00
Claudio Atzori	1902728f7e	Merge pull request '[ActionManagerFramework] documentation' (#347 ) from actionset_docs into beta Reviewed-on: D-Net/dnet-hadoop#347	2023-10-12 10:07:25 +02:00
Claudio Atzori	dda602fff7	[AMF] docs	2023-10-12 10:05:46 +02:00
Claudio Atzori	05ee7d8b09	[graph cleaning] avoid NPEs	2023-10-12 09:13:42 +02:00
Miriam Baglioni	8e9493fad9	mergin with branch beta	2023-10-11 18:18:09 +02:00
Miriam Baglioni	89184d5b4f	used the API instead of the IS for bulktagging and propagation for community through organization. Added a new propagation step for communities through projects. Still using the API and not the IS	2023-10-11 18:17:35 +02:00
Claudio Atzori	a460ebe215	[UnresolvedEntities] updated action name	2023-10-10 15:50:11 +02:00
Claudio Atzori	ecea58a41c	Merge pull request '[UnresolvedEntities] changing in the creation of the unresolved entities' (#346 ) from fos into beta Reviewed-on: D-Net/dnet-hadoop#346	2023-10-10 15:10:21 +02:00
Claudio Atzori	66064e99fe	Merge branch 'beta' into fos	2023-10-10 15:07:21 +02:00
Miriam Baglioni	a431b04814	leftover for the properties and removal of bipfinder	2023-10-10 12:53:57 +02:00
Claudio Atzori	ed9282ef2a	removed module dhp-stats-monitor-update	2023-10-10 09:52:03 +02:00
Miriam Baglioni	110ce4b40f	extend the fos model to include the level4 and the scores for level3 and level4. removed bip indicators from the instance	2023-10-10 09:46:40 +02:00
Claudio Atzori	204404b0e3	Merge branch 'beta' of https://code-repo.d4science.org/D-Net/dnet-hadoop into beta	2023-10-10 09:36:13 +02:00
Claudio Atzori	9a98f408b3	code formatting	2023-10-10 09:36:11 +02:00
Claudio Atzori	4e6fccf4f6	Merge pull request 'Beta stats wf updated' (#332 ) from antonis.lempesis/dnet-hadoop:beta into beta Reviewed-on: D-Net/dnet-hadoop#332	2023-10-10 09:35:32 +02:00
Miriam Baglioni	a3d01ccb24	refactoring	2023-10-09 14:52:17 +02:00
Miriam Baglioni	8448b9ebfb	mergin with branch beta	2023-10-09 14:27:23 +02:00
Miriam Baglioni	3d6be20989	changes to use the API instead of the IS the get the information for the communities to be used during bulktagging and context propagation	2023-10-09 14:26:33 +02:00
dimitrispie	17586f0ff8	Update step20-createMonitorDB.sql Add result_orcid table to monitor dbs	2023-10-09 14:21:31 +03:00
dimitrispie	489a082f04	Update step16-createIndicatorsTables.sql Change scripts for gold, hybrid, bronze indicators	2023-10-09 14:00:50 +03:00
Claudio Atzori	ef833840c3	[Doiboost] removed linkage to SFI unidentified project	2023-10-06 15:48:18 +02:00
Claudio Atzori	84a58802ab	[OC] using the common pid cleaning function	2023-10-06 14:48:05 +02:00
Claudio Atzori	46034630cf	[OC] compress the output actionset	2023-10-06 14:42:02 +02:00
Claudio Atzori	774e874d18	Merge pull request 'implemented relation to irish funder from a Json list' (#344 ) from irish_funder into beta Reviewed-on: D-Net/dnet-hadoop#344	2023-10-06 14:26:54 +02:00
Claudio Atzori	3bc44fbf1d	Merge branch 'beta' into irish_funder	2023-10-06 14:26:41 +02:00
Claudio Atzori	11153742c9	Merge pull request 'Extending the coverage of the peer non-unknown refereed instances' (#342 ) from peer_reviewed into beta Reviewed-on: D-Net/dnet-hadoop#342	2023-10-06 14:22:13 +02:00
Claudio Atzori	8108491722	Merge branch 'beta' into peer_reviewed	2023-10-06 14:21:52 +02:00
Giambattista Bloisi	2f3cf6d0e7	Fix cleaning of Pmid where parsing of numbers stopped at first not leading 0' character	2023-10-06 14:20:15 +02:00
Claudio Atzori	6856ab28ab	Merge pull request 'SWH_integration' (#343 ) from SWH_integration into beta Reviewed-on: D-Net/dnet-hadoop#343	2023-10-06 14:15:56 +02:00
Claudio Atzori	3c23d5f9bc	Merge branch 'beta' into SWH_integration	2023-10-06 14:15:38 +02:00
Claudio Atzori	858931ccb6	[SWH] compress the output actionset	2023-10-06 14:03:33 +02:00
Claudio Atzori	f759b18bca	[SWH] aligned parameter name	2023-10-06 13:43:20 +02:00
Claudio Atzori	eed9fe0902	code formatting	2023-10-06 12:31:17 +02:00
Claudio Atzori	7f27111b1f	Merge branch 'importpoci' into beta	2023-10-06 12:23:28 +02:00
Claudio Atzori	73c49b8d26	Merge branch 'beta' into SWH_integration	2023-10-06 12:21:51 +02:00
Sandro La Bruzzo	42a2dad975	implemented relation to irish funder from a Json list	2023-10-06 11:52:33 +02:00
Sandro La Bruzzo	13f332ce77	ignored jenv prop	2023-10-06 10:40:05 +02:00
Serafeim Chatzopoulos	1bb83b9188	Add prefix in SWH ID	2023-10-04 20:31:45 +03:00
Claudio Atzori	ee8a39e7d2	cleanup and refinements	2023-10-04 12:32:05 +02:00
Serafeim Chatzopoulos	e9f24df21c	Move SWH API Key from constants to workflow param	2023-10-03 20:57:57 +03:00
Serafeim Chatzopoulos	cae75fc75d	Add SWH in the collectedFrom field	2023-10-03 16:55:10 +03:00
Serafeim Chatzopoulos	b49a3ac9b2	Add actionsetsPath as a global WF param	2023-10-03 15:43:38 +03:00
Serafeim Chatzopoulos	24c43e0c60	Restructure workflow parameters	2023-10-03 15:11:58 +03:00
Serafeim Chatzopoulos	9f73d93e62	Add param for limiting repo Urls	2023-10-03 14:39:08 +03:00
Claudio Atzori	b446a9ed98	Merge branch 'beta' into peer_reviewed	2023-10-03 10:52:23 +02:00
Claudio Atzori	f344ad76d0	Merge pull request 'extended existing code to import of POCI from open citation' (#340 ) from importpoci into beta Reviewed-on: D-Net/dnet-hadoop#340	2023-10-03 10:52:11 +02:00
Claudio Atzori	5919e488dd	Merge branch 'beta' into importpoci	2023-10-03 10:43:53 +02:00
Serafeim Chatzopoulos	839a8524e7	Add action for creating actionsets	2023-10-02 23:50:38 +03:00
Claudio Atzori	c9a5ad6a02	extending the coverage of the peer non-unknown refereed instances	2023-10-02 16:28:42 +02:00
Miriam Baglioni	d7fccdc64b	fixed paths in wf to match the req of the pathname	2023-10-02 14:10:57 +02:00
Miriam Baglioni	9898470b0e	Addressing comments in D-Net/dnet-hadoop#340 \#issuecomment-10592	2023-10-02 12:54:16 +02:00
Giambattista Bloisi	c412dc162b	Fix bug in conversion from dedup json model to Spark Dataset of Rows: list of strings contained the json escaped representation of the value instead of the plain value, this caused instanceTypeMatch failures because of the leading and trailing double quotes	2023-10-02 11:34:51 +02:00
Claudio Atzori	5d09b7db8b	Merge pull request 'SparkPropagateRelation relations do not propagate deletedByInference and invisible' (#333 ) from consistency_keep_mergerels into beta Reviewed-on: D-Net/dnet-hadoop#333	2023-10-02 11:27:57 +02:00
Claudio Atzori	7b403a920f	Merge branch 'beta' into consistency_keep_mergerels	2023-10-02 11:26:00 +02:00
Claudio Atzori	dc86018a5f	Merge branch 'merge_entities_job' into beta	2023-10-02 11:24:48 +02:00
Giambattista Bloisi	3c47920c78	Use asScala to convert java List to Scala Sequence	2023-10-02 11:04:47 +02:00
Claudio Atzori	7f244d9a7a	code formatting	2023-10-02 11:04:36 +02:00
Giambattista Bloisi	e239b81740	Fix defect #8997 : GenerateEventsJob is generating huge amounts of logs because broker entity similarity calculation consistently failed	2023-10-02 11:04:18 +02:00
Miriam Baglioni	e84f5b5e64	extended existing codo to accomodate import of POCI from open citation	2023-10-02 09:25:16 +02:00
Serafeim Chatzopoulos	ab0d70691c	Add step for archiving repoUrls to SWH	2023-09-28 20:56:18 +03:00
Serafeim Chatzopoulos	ed9c81a0b7	Add steps to collect last visit data && archive not found repository URLs	2023-09-27 19:00:54 +03:00
Alessia Bardi	0935d7757c	Use v5 of the UNIBI Gold ISSN list in test	2023-09-20 15:41:35 +02:00
Alessia Bardi	cc7204a089	tests for d4science catalog	2023-09-20 15:38:32 +02:00
Sandro La Bruzzo	76476cdfb6	Added maven repo for dependencies that are not in maven central	2023-09-20 10:33:14 +02:00
dimitrispie	9ef971a146	Update step16-createIndicatorsTables.sql Fix int year for: indi_org_openess_year indi_org_fairness_year indi_org_findable_year	2023-09-19 14:25:42 +03:00
Serafeim Chatzopoulos	9d44418d38	Add collecting software code repository URLs	2023-09-14 18:43:25 +03:00
Serafeim Chatzopoulos	395a4af020	Run CC and RAM sequentieally in dhp-impact-indicators WF	2023-09-13 08:59:40 +02:00
Claudio Atzori	8a6892cc63	[graph dedup] consistency wf should not remove the relations while dispatching the entities	2023-09-12 21:27:05 +02:00
Claudio Atzori	4786aa0e09	added Archive ouverte UNIGE (ETHZ.UNIGENF, opendoar____::1400) to the Datacite hostedBy_map	2023-09-07 11:21:07 +02:00
dimitrispie	5f90cc11e9	Update step16-createIndicatorsTables.sql Fix indi_pub_bronze_oa	2023-09-06 14:14:38 +03:00
Giambattista Bloisi	2caaaec42d	Include SparkCleanRelation logic in SparkPropagateRelation SparkPropagateRelation includes merge relations Revised tests for SparkPropagateRelation	2023-09-04 11:33:20 +02:00
dimitrispie	964c2f553e	Changes in indicators step, monitor step - graduatedoctorates for observatory - result_apc_affiliations table - new indicators indi_is_funder_plan_s indi_funder_fairness indi_ris_fairness indi_funder_openess indi_ris_openess indi_funder_findable indi_ris_findable indi_is_project_result_after - cast year to int in composite indicators - new institutions -- Universidade Católica Portuguesa -- Iscte - Instituto Universitário de Lisboa -- Munster Technological University -- Cardiff University -- Leibniz Institute of Ecological Urban and Regional Development	2023-09-01 10:57:02 +03:00
Giambattista Bloisi	6cc7d8ca7b	GroupEntities and DispatchEntites are now merged in GroupEntitiesSparkJob	2023-08-30 10:43:31 +02:00
dimitrispie	be4856ef35	Update step15.sql	2023-07-17 15:33:58 +03:00
dimitrispie	163b2ee2a8	Changes 1. Monitor updates 2. Bug fixes during copy to impala cluster	2023-07-13 15:25:00 +03:00