Claudio Atzori
|
4d0c59669b
|
merged changes from beta
|
2024-01-26 16:08:54 +01:00 |
Claudio Atzori
|
9e8fc6aa88
|
[collection] increased logging from the oai-pmh metadata collection process
|
2024-01-26 09:17:20 +01:00 |
Claudio Atzori
|
3e96777cc4
|
[collection] increased logging from the oai-pmh metadata collection process
|
2024-01-23 15:21:03 +01:00 |
Claudio Atzori
|
1c6db320f4
|
[graph provision] obtain context info from the context API instead from the ISLookUp service
|
2024-01-22 15:53:17 +01:00 |
Claudio Atzori
|
1726f49790
|
code formatting
|
2023-12-15 10:37:02 +01:00 |
Claudio Atzori
|
e6086efc53
|
avoid NPEs in Vocabulary.getTermBySynonym
|
2023-12-03 13:33:20 +01:00 |
Claudio Atzori
|
5f1ed61c1f
|
merging from bulkTag branch
|
2023-11-03 12:51:37 +01:00 |
Miriam Baglioni
|
0097f4e64b
|
Removed Query community testing. Removed package from common related to the interaction with Zenodo since it was moved to the dump-project
|
2023-10-26 09:38:09 +02:00 |
Claudio Atzori
|
d28b7085f6
|
more NPE checks
|
2023-10-17 11:09:31 +02:00 |
Claudio Atzori
|
554551682d
|
[raw graph] adopting the new COAR based vocabularies for the resource typing
|
2023-10-11 16:09:19 +02:00 |
Serafeim Chatzopoulos
|
ab0d70691c
|
Add step for archiving repoUrls to SWH
|
2023-09-28 20:56:18 +03:00 |
Serafeim Chatzopoulos
|
ed9c81a0b7
|
Add steps to collect last visit data && archive not found repository URLs
|
2023-09-27 19:00:54 +03:00 |
Claudio Atzori
|
f3a85e224b
|
merged from branch beta the bulk tagging (single step, negative constraints), the cleanig worflow (single step, pid type based cleaning), instance level fulltext
|
2023-06-28 13:33:57 +02:00 |
Claudio Atzori
|
8a463cc3e8
|
fixed organization id created when mapping APC affiliations. Factored out ROR constants in dhp-common
|
2023-05-15 15:44:46 +02:00 |
Claudio Atzori
|
d02916ef82
|
code formatting
|
2023-05-02 11:05:37 +02:00 |
Miriam Baglioni
|
73f77575bd
|
[ZenodoApiClient] align with master version
|
2023-04-18 10:25:27 +02:00 |
Miriam Baglioni
|
087b5a7973
|
[ZenodiAPIClient] new version of the API to connect to Zenodo (change the http client
|
2023-04-17 18:59:22 +02:00 |
Miriam Baglioni
|
c6a7602b3e
|
refactoring after compilation
|
2023-04-06 14:45:01 +02:00 |
Miriam Baglioni
|
9a9cc6a1dd
|
changed the way the tar archive is build to support renaming in case we need to change .tt.gz into .json.gz
|
2023-04-04 11:40:58 +02:00 |
Miriam Baglioni
|
32870339f5
|
refactoring after compile
|
2023-02-13 13:06:48 +01:00 |
Sandro La Bruzzo
|
6c81a161d2
|
Merge remote-tracking branch 'origin/beta' into 8231-mdstore-synch-improve
|
2023-02-08 10:29:09 +01:00 |
Miriam Baglioni
|
b713132db7
|
[Cleaning] adding missing classes
|
2022-12-21 12:49:08 +01:00 |
Claudio Atzori
|
b8bafab8a0
|
[cleaning] improved vocabulary based mapping, specialization for the strict vocab cleaning
|
2022-12-12 14:43:03 +01:00 |
Sandro La Bruzzo
|
5a48a2fb18
|
implemented synch for single mdstore
|
2022-12-01 11:34:43 +01:00 |
Claudio Atzori
|
11695ba649
|
[graph cleaning] patch also the result's collectedfrom and hostedby datasource name according to the datasource master-duplicate mapping
|
2022-11-28 10:18:43 +01:00 |
Claudio Atzori
|
24ef301cc1
|
[graph cleaning] patch the result's collectedfrom and hostedby identifiers according to the datasource master-duplicate mapping
|
2022-11-28 09:54:18 +01:00 |
Claudio Atzori
|
adb526b0e1
|
Merge branch 'beta' into clean_subjects
|
2022-08-12 10:51:17 +02:00 |
Claudio Atzori
|
cb7c07c54e
|
[scholix] added step to create tar archive
|
2022-08-11 11:25:24 +02:00 |
Claudio Atzori
|
32cee1f619
|
WIP: cleaning of subjects
|
2022-08-05 12:32:08 +02:00 |
Claudio Atzori
|
b78889a0ce
|
WIP: cleaning of subjects
|
2022-08-05 09:11:37 +02:00 |
Claudio Atzori
|
1138b2ac8e
|
code formatting
|
2022-07-19 14:15:49 +02:00 |
Claudio Atzori
|
0cb1c70788
|
code formatting
|
2022-07-01 10:44:08 +02:00 |
Claudio Atzori
|
b295a40d9c
|
restored use of name_particles when parsing author names
|
2022-06-16 12:20:43 +02:00 |
Miriam Baglioni
|
ab8868bd3a
|
[ZENODO-API] changed to iterate in all the deposited products and not just the last ten
|
2022-06-08 17:03:15 +02:00 |
Miriam Baglioni
|
b7c2340952
|
[HostedByMap - DOIBoost] changed to use code moved to common since used also from hostedbymap now
|
2022-03-04 11:05:23 +01:00 |
Miriam Baglioni
|
56409d1281
|
[Dump] resolved conflicts with beta and merging
|
2021-12-14 15:03:45 +01:00 |
Miriam Baglioni
|
936578aaf1
|
Merge branch 'beta' of https://code-repo.d4science.org/D-Net/dnet-hadoop into beta
|
2021-12-13 15:01:47 +01:00 |
Claudio Atzori
|
e6e177dda0
|
vocabulary based cleaning considers also the term label when looking up for a synonym
|
2021-12-09 13:57:53 +01:00 |
Miriam Baglioni
|
96a7d46278
|
[Graph Dump] fixed tests
|
2021-12-06 15:06:32 +01:00 |
Miriam Baglioni
|
9fae872181
|
[Graph Dump] changed to mirror the changes in the model
|
2021-11-19 11:25:50 +01:00 |
Claudio Atzori
|
6b34ba737e
|
minor
|
2021-10-21 14:16:18 +02:00 |
Miriam Baglioni
|
c8321ad31a
|
merge with branch beta
|
2021-10-01 12:59:08 +02:00 |
Claudio Atzori
|
663b1556d7
|
manually integrating PR#140 #140
|
2021-09-15 16:40:25 +02:00 |
Claudio Atzori
|
3359f73fcf
|
cleanup & best practices
|
2021-08-13 12:00:42 +02:00 |
Miriam Baglioni
|
6e84b3951f
|
GetCSV refactoring - moving classes to dhp-common that have dependency with GetCSV class (that was located in graph-mapper)
|
2021-08-12 17:57:41 +02:00 |
Claudio Atzori
|
2ee21da43b
|
suggestions from SonarLint
|
2021-08-11 12:13:22 +02:00 |
Miriam Baglioni
|
d418c309f5
|
removed the part after part-x- in the file name generated by spark. It was too long and created problems while creating the tar entries
|
2021-07-13 17:11:49 +02:00 |
Claudio Atzori
|
9d725efdc1
|
reverted implementation of the mdstore client
|
2021-05-20 18:26:09 +02:00 |
Claudio Atzori
|
23b8883ab1
|
applied intellij code cleanup
|
2021-05-14 10:58:12 +02:00 |
Claudio Atzori
|
923d19ea8e
|
mdstore read lock/unlock when bulk copying records from mongodb to hdfs
|
2021-05-04 18:06:21 +02:00 |