dnet-hadoop

sabeel

Author	SHA1	Message	Date
Claudio Atzori	3f07390a58	WIP	2024-02-28 10:10:10 +01:00
Michele Artini	3268570b2c	mapping of project PIDs	2024-02-22 14:47:21 +01:00
Claudio Atzori	a63b091bae	Merge branch 'beta' into import_orps_fix	2024-02-15 15:01:56 +01:00
Claudio Atzori	d85d2df6ad	[graph raw] fixed mapping of the original resource type from the Datacite format	2024-02-09 10:20:20 +01:00
Giambattista Bloisi	b19643f6eb	Dedup aliases, created when a dedup in a previous build has been merged in a new dedup, need to be marked as "deletedbyinference", since they are "merged" in the new dedup	2024-02-08 15:34:59 +01:00
Claudio Atzori	38c9001147	fixed import of ORPs stored on HDFS in the internal graph format (e.g. Datacite)	2024-02-07 17:02:05 +01:00
Claudio Atzori	fd17c1f17c	[actiosets] fixed join type	2024-02-05 16:55:36 +02:00
Claudio Atzori	009dcf6aea	[actiosets] introduced support for the PromoteAction strategy	2024-02-05 16:43:40 +02:00
Claudio Atzori	42f5506306	[orcid enrichment] fixed directory cleanup before distcp	2024-02-05 09:45:36 +02:00
Alessia Bardi	f2a08d8cc2	test for Italian records from IRS repositories	2024-01-30 19:20:14 +01:00
Miriam Baglioni	a5995ab557	[orcid-enrichment] change the value of parameters.	2024-01-29 18:19:48 +01:00
Claudio Atzori	926903b06b	Merge branch 'beta' into stats_with_spark_sql	2024-01-29 09:11:45 +01:00
Giambattista Bloisi	078df0b4d1	Use SparkSQL in place of Hive for executing step16-createIndicatorsTables.sql of stats update wf	2024-01-26 21:56:55 +01:00
Claudio Atzori	ce3200263e	Merge branch 'beta' into crossref_missing_author_fix	2024-01-26 15:57:04 +01:00
Sandro La Bruzzo	e889808daa	Fixed problem on missing author in crossref Mapping	2024-01-26 12:19:04 +01:00
Antonis Lempesis	a7115cfa9e	max mem of joins (hive.mapjoin.followby.gby.localtask.max.memory.usage) now 80%, up from 55%.	2024-01-25 15:13:16 +01:00
Claudio Atzori	9b13c22e5d	[graph provision] retrieve all the context information by adding all=true to the requests issued to thr API	2024-01-23 15:36:08 +01:00
Claudio Atzori	f87f3a6483	[graph provision] updated param specification for the XML converter job	2024-01-23 08:54:37 +01:00
Claudio Atzori	6fd25cf549	code formatting	2024-01-23 08:47:12 +01:00
Claudio Atzori	f76852f385	Merge branch 'beta' into update_pivots_table	2024-01-22 16:37:22 +01:00
Claudio Atzori	1c6db320f4	[graph provision] obtain context info from the context API instead from the ISLookUp service	2024-01-22 15:53:17 +01:00
Claudio Atzori	2655eea5bc	[orcid enrichment] drop paths before copying the non-modifyed contents	2024-01-19 16:28:05 +01:00
Claudio Atzori	c6b3401596	increased shuffle partitions for publications in the country propagation workflow	2024-01-19 10:15:39 +01:00
Miriam Baglioni	bcc0a13981	[enrichment single step] adding <end> element in wf definition	2024-01-18 17:39:14 +01:00
Miriam Baglioni	6af536541d	[enrichment single step] moving parameter file in correct location	2024-01-18 15:35:40 +01:00
Miriam Baglioni	a12a3eb143	-	2024-01-18 15:18:10 +01:00
Miriam Baglioni	82e9e262ee	[enrichment single step] remove parameter from execution	2024-01-17 17:38:03 +01:00
Miriam Baglioni	67ce2d54be	[enrichment single step] refactoring to fix issues in disappeared result type	2024-01-17 16:50:00 +01:00
Miriam Baglioni	59eaccbd87	[enrichment single step] refactoring to fix issue in disappeared result type	2024-01-15 17:49:54 +01:00
Giambattista Bloisi	21a14fcd80	Reusable RunSQLSparkJob for executing SQL in Spark through Oozie Spark Actions Implements pivots table update oozie workflow	2024-01-15 10:18:14 +01:00
Miriam Baglioni	f612125939	fix issue on FoS integration. Removing the null values from FoS	2024-01-12 10:20:28 +01:00
Claudio Atzori	cb9e739484	Merge branch 'beta' into resource_types	2024-01-11 16:29:41 +01:00
Claudio Atzori	2753044d13	refined mapping for the extraction of the original resource type	2024-01-11 16:28:26 +01:00
Giambattista Bloisi	3c66e3bd7b	Create dedup record for "merged" pivots Do not create dedup records for group that have more than 20 different acceptance date	2024-01-10 22:59:52 +01:00
Giambattista Bloisi	10e135db1e	Use dedup_wf_002 in place of dedup_wf_001 to make explicit a different algorithm has been used to generate those kind of ids	2024-01-10 22:59:52 +01:00
Giambattista Bloisi	831cc1fdde	Generate "merged" dedup id relations also for records that are filtered out by the cut parameters	2024-01-10 22:59:52 +01:00
Giambattista Bloisi	1287315ffb	Do no longer use dedupId information from pivotHistory Database	2024-01-10 22:59:52 +01:00
Giambattista Bloisi	02636e802c	SparkCreateSimRels: - Create dedup blocks from the complete queue of records matching cluster key instead of truncating the results - Clean titles once before clustering and similarity comparisons - Added support for filtered fields in model - Added support for sorting List fields in model - Added new JSONListClustering and numAuthorsTitleSuffixPrefixChain clustering functions - Added new maxLengthMatch comparator function - Use reduced complexity Levenshtein with threshold in levensteinTitle - Use reduced complexity AuthorsMatch with threshold early-quit - Use incremental Connected Component to decrease comparisons in similarity match in BlockProcessor - Use new clusterings configuration in Dedup tests SparkWhitelistSimRels: use left semi join for clarity and performance SparkCreateMergeRels: - Use new connected component algorithm that converge faster than Spark GraphX provided algorithm - Refactored to use Windowing sorting rather than groupBy to reduce memory pressure - Use historical pivot table to generate singleton rels, merged rels and keep continuity with dedupIds used in the past - Comparator for pivot record selection now uses "tomorrow" as filler for missing or incorrect date instead of "2000-01-01" - Changed generation of ids of type dedup_wf_001 to avoid collisions DedupRecordFactory: use reduceGroups instead of mapGroups to decrease memory pressure	2024-01-10 22:59:52 +01:00
Miriam Baglioni	e711a05229	fixed conflicts	2024-01-10 11:03:42 +01:00
Miriam Baglioni	71d6f30711	Merge branch 'beta' of https://code-repo.d4science.org/D-Net/dnet-hadoop into beta	2024-01-10 10:59:58 +01:00
Miriam Baglioni	cb14470ba6	added properties file in the forlder for the workflow of result to organization from inst repo propagation. Changes the path in the classes implementing the propagation	2023-12-22 14:50:05 +01:00
Miriam Baglioni	9f966b59d4	added properties file in the forlder for the workflow of result to community from semrel propagation. Changes the path in the classes implementing the propagation	2023-12-22 14:11:47 +01:00
Miriam Baglioni	2f3b5a133d	added properties file in the forlder for the workflow of result to community from organization propagation. Changes the path in the classes implementing the propagation	2023-12-22 13:56:40 +01:00
Miriam Baglioni	2f7b9ad815	added properties file in the forlder for the workflow of project to result propagation. Changes the path in the classes implementing the propagation	2023-12-22 11:46:15 +01:00
Miriam Baglioni	f2352e8a78	changed in the classes the path for the property files for the propagation of community from project	2023-12-22 11:43:34 +01:00
Miriam Baglioni	009730b3d1	added properties file in the forlder for the workflow of orcid propagation. Changes the path in the classes implementing the propagationchanged the path to the parameter file in the class for entitytoorganization propagation	2023-12-22 11:42:09 +01:00
Miriam Baglioni	89f269c7f4	changed the path to the parameter file in the class for entitytoorganization propagation	2023-12-22 11:37:50 +01:00
Miriam Baglioni	b06aea0adf	adding the bulkTag parameter file in the folder for the oozie workflow for bulkTagging. Changes the path in the class	2023-12-22 11:35:37 +01:00
Miriam Baglioni	3afd4aa57b	adjustments for country propagation	2023-12-22 11:27:30 +01:00
Claudio Atzori	62104790ae	added metaresourcetype to the result hive DB view	2023-12-21 12:27:10 +01:00

1 2 3 4 5 ...

4086 Commits