Giambattista Bloisi
|
2f3cf6d0e7
|
Fix cleaning of Pmid where parsing of numbers stopped at first not leading 0' character
|
2023-10-06 14:20:15 +02:00 |
Giambattista Bloisi
|
2c235e82ad
|
Fix cleaning of Pmid where parsing of numbers stopped at first not leading 0' character
|
2023-10-06 12:35:54 +02:00 |
Claudio Atzori
|
eed9fe0902
|
code formatting
|
2023-10-06 12:31:17 +02:00 |
Claudio Atzori
|
73c49b8d26
|
Merge branch 'beta' into SWH_integration
|
2023-10-06 12:21:51 +02:00 |
Claudio Atzori
|
c9a5ad6a02
|
extending the coverage of the peer non-unknown refereed instances
|
2023-10-02 16:28:42 +02:00 |
Serafeim Chatzopoulos
|
ab0d70691c
|
Add step for archiving repoUrls to SWH
|
2023-09-28 20:56:18 +03:00 |
Serafeim Chatzopoulos
|
ed9c81a0b7
|
Add steps to collect last visit data && archive not found repository URLs
|
2023-09-27 19:00:54 +03:00 |
Claudio Atzori
|
8a6892cc63
|
[graph dedup] consistency wf should not remove the relations while dispatching the entities
|
2023-09-12 21:27:05 +02:00 |
Claudio Atzori
|
dc80ab14d3
|
[graph dedup] consistency wf should not remove the relations while dispatching the entities
|
2023-09-12 14:34:28 +02:00 |
Giambattista Bloisi
|
6cc7d8ca7b
|
GroupEntities and DispatchEntites are now merged in GroupEntitiesSparkJob
|
2023-08-30 10:43:31 +02:00 |
Claudio Atzori
|
bf35280ea6
|
code formatting
|
2023-08-29 11:11:00 +02:00 |
Giambattista Bloisi
|
95cd2b9b1e
|
Make filterInvisible a mandatory parameter of DispathEntitiesSparkJob
Make filterInvisible a mandatory parameter of both dedup/consistency and graph/group oozie workflows
|
2023-08-10 11:53:48 +02:00 |
Giambattista Bloisi
|
fab9920271
|
DispatchEntitiesSparkJob: manage all entity types together, support filtering by dataInfo.invisible flag
|
2023-08-09 15:41:43 +02:00 |
Miriam Baglioni
|
599828ce35
|
Merge branch 'master' of https://code-repo.d4science.org/D-Net/dnet-hadoop
|
2023-08-09 13:07:13 +02:00 |
Miriam Baglioni
|
c25ac21e5e
|
Merge pull request 'graph cleaning, suggestions from ticket 8898' (#325) from cleaning_8898 into beta
Reviewed-on: #325
|
2023-08-08 11:14:19 +02:00 |
Claudio Atzori
|
0bc74e2000
|
code formatting
|
2023-08-02 11:52:10 +02:00 |
Claudio Atzori
|
7180911ded
|
[graph cleaning] fixed regex behaviour for cleaning ROR and GRID identifiers, added tests
|
2023-08-02 11:44:14 +02:00 |
Claudio Atzori
|
b9dddbfe54
|
rule out records with NULL dataInfo, except for Relations
|
2023-07-31 17:53:54 +02:00 |
Claudio Atzori
|
da1727f93f
|
rule out records with NULL dataInfo, except for Relations
|
2023-07-31 17:52:56 +02:00 |
Claudio Atzori
|
11ffb9bd68
|
rule out records with NULL dataInfo
|
2023-07-31 12:35:33 +02:00 |
Claudio Atzori
|
ccac6a7f75
|
rule out records with NULL dataInfo
|
2023-07-31 12:35:05 +02:00 |
Claudio Atzori
|
d512df8612
|
code formatting
|
2023-07-26 09:14:08 +02:00 |
Claudio Atzori
|
d8435a6512
|
inverted condition
|
2023-07-25 17:39:57 +02:00 |
Claudio Atzori
|
59764145bb
|
cherry picked & fixed commit 270df939c4
|
2023-07-25 17:39:00 +02:00 |
Claudio Atzori
|
270df939c4
|
partial implementation of the suggestions from https://support.openaire.eu/issues/8898
|
2023-07-25 17:29:50 +02:00 |
Claudio Atzori
|
c754397a19
|
Merge branch 'beta' into pid_cleaning
|
2023-07-24 10:49:31 +02:00 |
Giambattista Bloisi
|
38dfebfbe6
|
Disable MdStoreClientTest test as it requires a local mongodb running and it does not perform any assertions
|
2023-07-19 14:18:56 +02:00 |
Miriam Baglioni
|
9e8e39f78a
|
-
|
2023-07-19 11:35:58 +02:00 |
Claudio Atzori
|
f3a85e224b
|
merged from branch beta the bulk tagging (single step, negative constraints), the cleanig worflow (single step, pid type based cleaning), instance level fulltext
|
2023-06-28 13:33:57 +02:00 |
Sandro La Bruzzo
|
9910ce06ae
|
added to CreateSimRel the feature to write time log
|
2023-06-28 11:38:16 +02:00 |
Sandro La Bruzzo
|
b195da3a83
|
Added utility to write time logs during the deduplication phase
|
2023-06-28 11:20:09 +02:00 |
Claudio Atzori
|
0f5a819f44
|
[graph cleaning] fixed regex behaviour for cleaning ROR and GRID identifiers, added tests
|
2023-06-23 16:10:49 +02:00 |
Miriam Baglioni
|
e4b27182d0
|
[master] refactoring
|
2023-06-21 11:15:53 +02:00 |
Claudio Atzori
|
1d33074fd1
|
WIP: pid cleaning
|
2023-06-09 16:47:25 +02:00 |
Miriam Baglioni
|
d9506035e4
|
[ZenodoApi] gone back to okhttp3 to send the payload.
|
2023-06-09 12:05:02 +02:00 |
Claudio Atzori
|
8a463cc3e8
|
fixed organization id created when mapping APC affiliations. Factored out ROR constants in dhp-common
|
2023-05-15 15:44:46 +02:00 |
Claudio Atzori
|
d02916ef82
|
code formatting
|
2023-05-02 11:05:37 +02:00 |
Claudio Atzori
|
851f664bd9
|
Merge branch 'beta' into graph_cleaning_refactoring
|
2023-05-02 09:55:40 +02:00 |
Miriam Baglioni
|
9fc8ebe98b
|
refactoring
|
2023-04-19 09:32:13 +02:00 |
Miriam Baglioni
|
73f77575bd
|
[ZenodoApiClient] align with master version
|
2023-04-18 10:25:27 +02:00 |
Miriam Baglioni
|
24c41806ac
|
[ZenodoApiClienttest] change test to mirror change in the omplementation
|
2023-04-18 09:08:09 +02:00 |
Miriam Baglioni
|
087b5a7973
|
[ZenodiAPIClient] new version of the API to connect to Zenodo (change the http client
|
2023-04-17 18:59:22 +02:00 |
Miriam Baglioni
|
c6a7602b3e
|
refactoring after compilation
|
2023-04-06 14:45:01 +02:00 |
Claudio Atzori
|
2a6ba29b64
|
[graph cleaning] unit tests & cleanup
|
2023-04-04 12:34:51 +02:00 |
Miriam Baglioni
|
9a9cc6a1dd
|
changed the way the tar archive is build to support renaming in case we need to change .tt.gz into .json.gz
|
2023-04-04 11:40:58 +02:00 |
Claudio Atzori
|
6d3d18d8b5
|
[graph cleaning] WIP: refactoring of the cleaning stages
|
2023-03-16 17:23:36 +01:00 |
Miriam Baglioni
|
32870339f5
|
refactoring after compile
|
2023-02-13 13:06:48 +01:00 |
Sandro La Bruzzo
|
0b9819f1ab
|
Code formatted
|
2023-02-08 10:32:33 +01:00 |
Sandro La Bruzzo
|
6c81a161d2
|
Merge remote-tracking branch 'origin/beta' into 8231-mdstore-synch-improve
|
2023-02-08 10:29:09 +01:00 |
Miriam Baglioni
|
b713132db7
|
[Cleaning] adding missing classes
|
2022-12-21 12:49:08 +01:00 |
Claudio Atzori
|
9cf0a98699
|
[cleaning] set the common subject classid/name
|
2022-12-20 10:17:33 +01:00 |
Claudio Atzori
|
b8bafab8a0
|
[cleaning] improved vocabulary based mapping, specialization for the strict vocab cleaning
|
2022-12-12 14:43:03 +01:00 |
Sandro La Bruzzo
|
5a48a2fb18
|
implemented synch for single mdstore
|
2022-12-01 11:34:43 +01:00 |
Claudio Atzori
|
11695ba649
|
[graph cleaning] patch also the result's collectedfrom and hostedby datasource name according to the datasource master-duplicate mapping
|
2022-11-28 10:18:43 +01:00 |
Claudio Atzori
|
24ef301cc1
|
[graph cleaning] patch the result's collectedfrom and hostedby identifiers according to the datasource master-duplicate mapping
|
2022-11-28 09:54:18 +01:00 |
Claudio Atzori
|
b47aaf4dd1
|
[cleaning] subjects declared as belonging to specific vocabularies whose values are not found in the vocab are set to type keyword
|
2022-10-13 11:23:43 +02:00 |
Claudio Atzori
|
b7c387c21f
|
cleaning of subjects: avoid duplicated subjects, prioritise collected vs inferred or other sources
|
2022-08-12 15:09:16 +02:00 |
Claudio Atzori
|
adb526b0e1
|
Merge branch 'beta' into clean_subjects
|
2022-08-12 10:51:17 +02:00 |
Claudio Atzori
|
cb7c07c54e
|
[scholix] added step to create tar archive
|
2022-08-11 11:25:24 +02:00 |
Claudio Atzori
|
3418ce50ac
|
cleaning of subjects: perform the cleaning when the given value is equivalent to one of the terms in the vocabulary
|
2022-08-08 12:48:47 +02:00 |
Claudio Atzori
|
32cee1f619
|
WIP: cleaning of subjects
|
2022-08-05 12:32:08 +02:00 |
Claudio Atzori
|
b78889a0ce
|
WIP: cleaning of subjects
|
2022-08-05 09:11:37 +02:00 |
Claudio Atzori
|
27a91841e7
|
WIP: cleaning of subjects
|
2022-08-04 11:39:39 +02:00 |
Claudio Atzori
|
09ccc7b472
|
Merge branch 'beta' into project_organization_contribution
|
2022-07-28 09:49:59 +02:00 |
Claudio Atzori
|
1138b2ac8e
|
code formatting
|
2022-07-19 14:15:49 +02:00 |
Claudio Atzori
|
0cb1c70788
|
code formatting
|
2022-07-01 10:44:08 +02:00 |
Claudio Atzori
|
7da24c1dec
|
added more logging
|
2022-06-28 13:47:49 +02:00 |
Claudio Atzori
|
a8773af0cb
|
Merge branch 'beta' into project_organization_contribution
|
2022-06-27 09:37:40 +02:00 |
Claudio Atzori
|
316b0fd73c
|
added 'von' to the name particles file
|
2022-06-27 09:36:51 +02:00 |
Claudio Atzori
|
5130eac247
|
mapping by participant project contribution
|
2022-06-24 17:16:42 +02:00 |
Claudio Atzori
|
b295a40d9c
|
restored use of name_particles when parsing author names
|
2022-06-16 12:20:43 +02:00 |
Miriam Baglioni
|
ab8868bd3a
|
[ZENODO-API] changed to iterate in all the deposited products and not just the last ten
|
2022-06-08 17:03:15 +02:00 |
Claudio Atzori
|
da611cfbbd
|
[eosc_services] resolved merge conflicts
|
2022-05-03 13:37:15 +02:00 |
Claudio Atzori
|
f5f532d134
|
EOSC Services - ongoing update
|
2022-04-29 12:25:24 +02:00 |
Miriam Baglioni
|
b61efd613b
|
[Measures] addressed comments in the PR
|
2022-04-21 12:09:37 +02:00 |
Miriam Baglioni
|
c304657d91
|
[Measures] put the logic in common, no need to change the schema
|
2022-04-21 11:27:26 +02:00 |
Miriam Baglioni
|
b7c2340952
|
[HostedByMap - DOIBoost] changed to use code moved to common since used also from hostedbymap now
|
2022-03-04 11:05:23 +01:00 |
Alessia Bardi
|
6158170334
|
testing delegated authority and bumped dep to schemas
|
2022-02-11 18:05:18 +01:00 |
Claudio Atzori
|
db299dd8ab
|
fixed typo
|
2022-01-27 16:24:06 +01:00 |
Claudio Atzori
|
c42623f006
|
added NPE checks
|
2022-01-21 14:30:09 +01:00 |
Claudio Atzori
|
391aa1373b
|
added unit test
|
2022-01-19 17:13:21 +01:00 |
Claudio Atzori
|
62f135262e
|
code formatting
|
2022-01-19 12:30:52 +01:00 |
Claudio Atzori
|
44a937f4ed
|
factored out entity grouping implementation, extended to consider results from delegated authorities rather than identical records from other sources
|
2022-01-19 12:24:52 +01:00 |
Miriam Baglioni
|
42e8f76778
|
[GraphCleaning] change the return value in the filtering function to avoid to lose the APC entities
|
2022-01-13 16:06:43 +01:00 |
Claudio Atzori
|
4f212652ca
|
scalafmt: code formatting
|
2022-01-11 16:57:48 +01:00 |
Miriam Baglioni
|
be0acccf42
|
Merge branch 'beta' into dump
|
2021-12-22 12:39:57 +01:00 |
Sandro La Bruzzo
|
3920d68992
|
Fixed workflow generation of delta in datacite
|
2021-12-21 11:41:49 +01:00 |
Sandro La Bruzzo
|
b881ee5ef8
|
[scholexplorer]
- implemented generation of scholix of delta update of datacite
|
2021-12-15 11:25:32 +01:00 |
Miriam Baglioni
|
56409d1281
|
[Dump] resolved conflicts with beta and merging
|
2021-12-14 15:03:45 +01:00 |
Miriam Baglioni
|
a3592b463a
|
Merge branch 'beta' of https://code-repo.d4science.org/D-Net/dnet-hadoop into beta
|
2021-12-14 14:58:26 +01:00 |
Claudio Atzori
|
aff3ddc8d2
|
added cleaning for the format field, removing carrige return and tab characters
|
2021-12-14 11:41:46 +01:00 |
Miriam Baglioni
|
936578aaf1
|
Merge branch 'beta' of https://code-repo.d4science.org/D-Net/dnet-hadoop into beta
|
2021-12-13 15:01:47 +01:00 |
Claudio Atzori
|
41c70c607d
|
cleaning workflow assigns the proper default instance type when a value could not be cleaned using the vocabularies
|
2021-12-09 16:44:28 +01:00 |
Claudio Atzori
|
e6e177dda0
|
vocabulary based cleaning considers also the term label when looking up for a synonym
|
2021-12-09 13:57:53 +01:00 |
Miriam Baglioni
|
b113586207
|
resolved conflicts
|
2021-12-07 10:16:14 +01:00 |
Sandro La Bruzzo
|
5d51b3dd4a
|
Merge pull request 'scala_refactor' (#169) from scala_refactor into beta
Reviewed-on: #169
|
2021-12-06 15:33:44 +01:00 |
Miriam Baglioni
|
96a7d46278
|
[Graph Dump] fixed tests
|
2021-12-06 15:06:32 +01:00 |
Sandro La Bruzzo
|
81bf604059
|
[scala-refactor] Module dhp-common:
Moved all scala source into src/main/scala and src/test/scala
|
2021-12-06 11:29:24 +01:00 |
Claudio Atzori
|
9132727793
|
fixed date cleaning test
|
2021-12-06 10:54:05 +01:00 |
Claudio Atzori
|
863a2f9db3
|
avoid to filter OAF records defined as invisible = true
|
2021-12-03 09:08:12 +01:00 |
Miriam Baglioni
|
8905a39bf3
|
mergin with branch beta
|
2021-12-02 13:17:29 +01:00 |
Sandro La Bruzzo
|
1e1f5e4fe0
|
minor fix
|
2021-11-25 13:03:17 +01:00 |
Sandro La Bruzzo
|
2164a2a889
|
Datacite: Code Refactor generated a general SparkApplication Scala where all the spark scala have to inherit
Commented a little the Datacite transformation code
|
2021-11-25 10:54:13 +01:00 |
Miriam Baglioni
|
9fae872181
|
[Graph Dump] changed to mirror the changes in the model
|
2021-11-19 11:25:50 +01:00 |
Claudio Atzori
|
82a4e4efae
|
[cleaning wf] fixed methodology to rule out invalid result titles, based on https://support.openaire.eu/issues/7206
|
2021-11-17 14:17:22 +01:00 |
Claudio Atzori
|
49f897ef29
|
[cleaning wf] fixed regex used to spot garbage in result titles; adjusted threshold for filtering titles
|
2021-11-16 15:24:23 +01:00 |
Sandro La Bruzzo
|
aafdffa6b3
|
resolved conflict
|
2021-10-26 09:45:46 +02:00 |
Sandro La Bruzzo
|
034304b33a
|
conflict resolved on merge
|
2021-10-26 09:40:47 +02:00 |
Claudio Atzori
|
6b34ba737e
|
minor
|
2021-10-21 14:16:18 +02:00 |
Sandro La Bruzzo
|
ae4e99a471
|
Adapted workflow of resolution of PID to work into OpenAIRE data workflow
- Added relations in both verse on all Scholexplorer datasources
|
2021-10-20 17:12:16 +02:00 |
Miriam Baglioni
|
c8321ad31a
|
merge with branch beta
|
2021-10-01 12:59:08 +02:00 |
Claudio Atzori
|
663b1556d7
|
manually integrating PR#140 #140
|
2021-09-15 16:40:25 +02:00 |
Claudio Atzori
|
baed5e3337
|
test classes moved in specific components
|
2021-08-13 12:14:47 +02:00 |
Claudio Atzori
|
3359f73fcf
|
cleanup & best practices
|
2021-08-13 12:00:42 +02:00 |
Miriam Baglioni
|
58f241f4a2
|
GetCSV refactoring - changed due to change of input resource
|
2021-08-13 10:04:44 +02:00 |
Miriam Baglioni
|
f3d575f749
|
GetCSV refactoring - changed due to changes in input resource
|
2021-08-13 10:03:57 +02:00 |
Miriam Baglioni
|
a5f6edfa6c
|
GetCSV refactoring - changed to mirror the original model class
|
2021-08-13 09:30:03 +02:00 |
Miriam Baglioni
|
733bcaecf6
|
GetCSV refactoring - added test class (all the tests are disabled since they refer to remote resource)
|
2021-08-12 17:58:52 +02:00 |
Miriam Baglioni
|
bfe8f5335c
|
GetCSV refactoring - copied model classes in test path
|
2021-08-12 17:58:14 +02:00 |
Miriam Baglioni
|
6e84b3951f
|
GetCSV refactoring - moving classes to dhp-common that have dependency with GetCSV class (that was located in graph-mapper)
|
2021-08-12 17:57:41 +02:00 |
Miriam Baglioni
|
f9b6b45d85
|
reverting
|
2021-08-11 17:04:48 +02:00 |
Miriam Baglioni
|
8da3a25cf6
|
merging with branch beta
|
2021-08-11 15:55:34 +02:00 |
Claudio Atzori
|
2ee21da43b
|
suggestions from SonarLint
|
2021-08-11 12:13:22 +02:00 |
Miriam Baglioni
|
6bd1eca7e0
|
merge branch with beta
|
2021-08-05 15:23:32 +02:00 |
Miriam Baglioni
|
ee13da9258
|
merge branch with master
|
2021-08-05 11:34:20 +02:00 |
Miriam Baglioni
|
1d6ac3715b
|
merge branch with beta
|
2021-07-30 11:58:29 +02:00 |
Claudio Atzori
|
a9961a1835
|
[cleaning] title cleaning based on the me.xuender:unidecode library
|
2021-07-28 16:36:33 +02:00 |
Claudio Atzori
|
6dddad86ee
|
[cleaning] title cleaning based on the me.xuender:unidecode library
|
2021-07-28 16:21:29 +02:00 |
Miriam Baglioni
|
74f801b689
|
mergin with branch beta
|
2021-07-27 13:18:31 +02:00 |
Miriam Baglioni
|
35e395eae8
|
merge with master
|
2021-07-27 12:34:59 +02:00 |
Miriam Baglioni
|
eb07f7f40f
|
Hosted By Map
|
2021-07-27 12:27:26 +02:00 |
Claudio Atzori
|
bc835d2024
|
[cleaning] fixed filtering function for missing titles
|
2021-07-23 11:56:13 +02:00 |
Claudio Atzori
|
ffdb2a3ea3
|
[cleaning] fixed filtering function for missing titles
|
2021-07-23 11:55:55 +02:00 |
Sandro La Bruzzo
|
62ae36a3d2
|
fixed NPE
|
2021-07-22 15:41:38 +02:00 |
Sandro La Bruzzo
|
d94565862a
|
fixed NPE
|
2021-07-21 21:23:11 +02:00 |
Sandro La Bruzzo
|
31d2d6d41e
|
Scholexplorer: introduction of dedup openaire
|
2021-07-21 18:09:32 +02:00 |
Miriam Baglioni
|
d418c309f5
|
removed the part after part-x- in the file name generated by spark. It was too long and created problems while creating the tar entries
|
2021-07-13 17:11:49 +02:00 |
Sandro La Bruzzo
|
ad50415167
|
Merge remote-tracking branch 'origin/stable_ids' into stable_id_scholexplorer
|
2021-06-24 17:20:50 +02:00 |
Claudio Atzori
|
67afd06cd1
|
[cleaning] cleaning instance.pid and instance.alternateidentifier using the same procedure used to clean result.pid
|
2021-06-24 12:10:17 +02:00 |
Sandro La Bruzzo
|
cc0f2b11fb
|
Implemented mapping from pubmed baseline to OAF
|
2021-06-16 14:56:24 +02:00 |
Claudio Atzori
|
2039bb9f5f
|
orcid / orcid_pending cleaning backported from master branch
|
2021-06-14 09:40:50 +02:00 |
Claudio Atzori
|
a900bfb874
|
delegating the date parsing to https://github.com/sisyphsu/dateparser
|
2021-06-11 16:53:01 +02:00 |
Claudio Atzori
|
eb6acfbabc
|
[cleaning] removing non parsable relation.validationDate(s)
|
2021-05-28 10:50:44 +02:00 |
Claudio Atzori
|
9d725efdc1
|
reverted implementation of the mdstore client
|
2021-05-20 18:26:09 +02:00 |
Claudio Atzori
|
23b8883ab1
|
applied intellij code cleanup
|
2021-05-14 10:58:12 +02:00 |
Claudio Atzori
|
d4c3476152
|
mapping datasource.journal only when an issn is available, null otherwhise
|
2021-05-11 11:08:54 +02:00 |
Claudio Atzori
|
d1cbee8413
|
imported methods from CleaningFunctions, defined in GraphCleaningFunctions
|
2021-05-10 16:43:39 +02:00 |
Claudio Atzori
|
3797543600
|
MDStoreManager model classes moved in dhp-schemas
|
2021-05-10 14:32:05 +02:00 |
Claudio Atzori
|
b1785ba77c
|
alternative way to set timeouts for the ISLookup client
|
2021-05-05 11:23:46 +02:00 |
Claudio Atzori
|
923d19ea8e
|
mdstore read lock/unlock when bulk copying records from mongodb to hdfs
|
2021-05-04 18:06:21 +02:00 |