Miriam Baglioni
|
d6895f0387
|
Merge branch 'beta' of https://code-repo.d4science.org/D-Net/dnet-hadoop into beta
|
2023-01-09 17:28:38 +01:00 |
dimitrispie
|
becb242c17
|
Monitor DB only Workflow
|
2023-01-04 16:50:29 +02:00 |
dimitrispie
|
dcb958e146
|
Changes to execute the stats wf only in hive
|
2023-01-04 11:39:01 +02:00 |
dimitrispie
|
592013d5dd
|
Added more steps in decision node
|
2022-12-23 09:43:16 +02:00 |
dimitrispie
|
2a4bf32d4c
|
Merge branch 'hive' of https://code-repo.d4science.org/antonis.lempesis/dnet-hadoop into hive
# Conflicts:
# dhp-workflows/dhp-stats-update/src/main/resources/eu/dnetlib/dhp/oa/graph/stats/oozie_app/scripts/step10.sql
# dhp-workflows/dhp-stats-update/src/main/resources/eu/dnetlib/dhp/oa/graph/stats/oozie_app/scripts/step13.sql
# dhp-workflows/dhp-stats-update/src/main/resources/eu/dnetlib/dhp/oa/graph/stats/oozie_app/scripts/step14.sql
# dhp-workflows/dhp-stats-update/src/main/resources/eu/dnetlib/dhp/oa/graph/stats/oozie_app/scripts/step16_1-definitions.sql
# dhp-workflows/dhp-stats-update/src/main/resources/eu/dnetlib/dhp/oa/graph/stats/oozie_app/scripts/step7.sql
|
2022-12-22 10:22:46 +02:00 |
dimitrispie
|
6449ff4207
|
1. Added a decision node to enables the workflow to make a selection on the execution path to follow
2. Added new organization
3. Added 5 new tables from Eurostast
|
2022-12-22 10:18:21 +02:00 |
Miriam Baglioni
|
8893389895
|
Merge branch 'beta' of https://code-repo.d4science.org/D-Net/dnet-hadoop into beta
|
2022-12-21 12:42:27 +01:00 |
Antonis Lempesis
|
c8309fe18e
|
addded command line params to allow hive actions to run
|
2022-12-21 12:41:33 +02:00 |
Antonis Lempesis
|
028873cc51
|
added new hive opts
|
2022-12-21 12:41:33 +02:00 |
Antonis Lempesis
|
1ddea4f442
|
removed 'stored as parquet' from views..
|
2022-12-21 12:41:33 +02:00 |
Antonis Lempesis
|
2754c3dd62
|
moving data to impala cluster and creating shadow databases there
|
2022-12-21 12:41:29 +02:00 |
Antonis Lempesis
|
778a1a724f
|
finished migration to hive only
|
2022-12-21 12:41:25 +02:00 |
Antonis Lempesis
|
e84dd5fe26
|
first
|
2022-12-21 12:41:23 +02:00 |
Sandro La Bruzzo
|
3c9826f186
|
updated lines function to it's implementation linesWithSeparators.map(l => l.stripLineEnd) in this way we force scala plugin compiler to consider this pipeline scala code and not java.string.lines() pipeline
|
2022-12-21 11:21:17 +01:00 |
Claudio Atzori
|
6aa91204a5
|
[orcid propagation] skip empty directories
|
2022-12-20 14:15:46 +01:00 |
Miriam Baglioni
|
6674cccb94
|
[BulkTag] description of parameters more comprehensive for those who do not implement it
|
2022-12-16 15:33:20 +01:00 |
Miriam Baglioni
|
f37113a941
|
[BulkTag] moving xquery to get community configuration in dedicated file
|
2022-12-16 15:32:26 +01:00 |
Miriam Baglioni
|
8685eaa706
|
[Clean Country] added test to verify remove of country
|
2022-12-16 15:31:25 +01:00 |
Miriam Baglioni
|
dc0ec88a58
|
Merge branch 'beta' of https://code-repo.d4science.org/D-Net/dnet-hadoop into beta
|
2022-12-16 13:18:32 +01:00 |
Miriam Baglioni
|
d791840b82
|
[Clean Country] added test to verify remove of country:
|
2022-12-16 13:18:29 +01:00 |
Claudio Atzori
|
7b80b24f82
|
[cleaning] country cleaning must use both PID and AlternateIdentifier fields
|
2022-12-15 14:49:04 +01:00 |
Claudio Atzori
|
b8bafab8a0
|
[cleaning] improved vocabulary based mapping, specialization for the strict vocab cleaning
|
2022-12-12 14:43:03 +01:00 |
Sandro La Bruzzo
|
5e4866d033
|
implemented synch for single mdstore
|
2022-12-12 11:29:46 +01:00 |
Claudio Atzori
|
c18b8048c3
|
[cleaning] avoid NPE
|
2022-12-10 11:41:38 +01:00 |
Claudio Atzori
|
8b44afe5e5
|
[cleaning] avoid NPE
|
2022-12-09 15:44:57 +01:00 |
Claudio Atzori
|
389dd25430
|
[cleaning] avoid NPE
|
2022-12-08 18:40:48 +01:00 |
Claudio Atzori
|
730228d73d
|
[cleaning] align wf parameter names in test
|
2022-12-08 18:40:22 +01:00 |
Claudio Atzori
|
2094fa6db0
|
[cleaning] align wf parameter names
|
2022-12-08 17:22:26 +01:00 |
Miriam Baglioni
|
a485a94956
|
[Cleaning] fixed parameter name in property file
|
2022-12-08 16:59:34 +01:00 |
Miriam Baglioni
|
3d99b78d94
|
[Cleaning] fixed error in parameter (workingPath to workingDir)
|
2022-12-08 10:25:02 +01:00 |
Claudio Atzori
|
1b8488976b
|
code formatting
|
2022-12-07 10:45:38 +01:00 |
Claudio Atzori
|
cd1b58483e
|
[bulk tag] fixed Community configuration parsing to void NPE
|
2022-12-07 10:39:00 +01:00 |
Claudio Atzori
|
062abfd669
|
fixed NPE, removed unused stuff
|
2022-12-06 12:04:00 +01:00 |
dimitrispie
|
2a52a42169
|
Added 4 institutions:
-University of Modena and Reggio Emilia
-Bilkent University
-Saints Cyril and Methodius University of Skopje
-University of Milan
|
2022-12-06 10:10:21 +02:00 |
Claudio Atzori
|
8248da40d9
|
Merge branch 'beta' into graph_cleaning
|
2022-12-02 14:49:00 +01:00 |
Claudio Atzori
|
ddf065756f
|
Merge pull request 'Two organizations are added for monitor' (#258) from antonis.lempesis/dnet-hadoop:beta into beta
Reviewed-on: D-Net/dnet-hadoop#258
|
2022-12-02 14:45:27 +01:00 |
Sandro La Bruzzo
|
5a48a2fb18
|
implemented synch for single mdstore
|
2022-12-01 11:34:43 +01:00 |
Claudio Atzori
|
a38116546d
|
Merge branch 'beta' into deduptesting
|
2022-11-30 11:27:29 +01:00 |
Miriam Baglioni
|
ce020f2c83
|
[EOSC FUTURE] added resources and test for review
|
2022-11-30 09:57:30 +01:00 |
Miriam Baglioni
|
bb0ddc1c44
|
[BulkTag] adding verb starts_with
|
2022-11-30 09:56:24 +01:00 |
Claudio Atzori
|
8e3edba318
|
[graph cleaning] testing the collectedfron and hostedby patch procedure
|
2022-11-29 16:07:09 +01:00 |
Claudio Atzori
|
58c05731f9
|
[graph cleaning] WIP: testing the collectedfron and hostedby patch procedure
|
2022-11-29 11:21:51 +01:00 |
Miriam Baglioni
|
9c70c5dbd6
|
[Bulk Tag horizontal] added new path in definition of constraint (to recognize fos subjects) - changed test and resource class to test this new aspect
|
2022-11-28 14:51:20 +01:00 |
Miriam Baglioni
|
0628df7a3a
|
resolving conflicts
|
2022-11-28 10:44:56 +01:00 |
Claudio Atzori
|
11695ba649
|
[graph cleaning] patch also the result's collectedfrom and hostedby datasource name according to the datasource master-duplicate mapping
|
2022-11-28 10:18:43 +01:00 |
Claudio Atzori
|
6082d235d3
|
Merge branch 'beta' of https://code-repo.d4science.org/D-Net/dnet-hadoop into graph_cleaning
|
2022-11-28 09:54:48 +01:00 |
Claudio Atzori
|
24ef301cc1
|
[graph cleaning] patch the result's collectedfrom and hostedby identifiers according to the datasource master-duplicate mapping
|
2022-11-28 09:54:18 +01:00 |
Alessia Bardi
|
90c8f9cb61
|
tests for EOSC Future
|
2022-11-23 12:18:44 +01:00 |
Miriam Baglioni
|
0e3edc5018
|
[Bulk Tag] fixed issue in verb name
|
2022-11-23 11:26:36 +01:00 |
Claudio Atzori
|
a79c47522d
|
updated ORCID datasource identifier
|
2022-11-23 10:17:49 +01:00 |
Alessia Bardi
|
2832117f23
|
added eoscifguidelines in test
|
2022-11-22 18:01:12 +01:00 |
Alessia Bardi
|
3c08269a4d
|
Merge branch 'beta' of https://code-repo.d4science.org/D-Net/dnet-hadoop into beta
|
2022-11-22 17:31:00 +01:00 |
Alessia Bardi
|
2687fc9f73
|
tests for EOSC Future review - ROhub
|
2022-11-22 17:30:56 +01:00 |
Claudio Atzori
|
1d5143b0b6
|
Merge branch 'beta' into deduptesting
|
2022-11-22 10:21:30 +01:00 |
Claudio Atzori
|
0aa725083f
|
extended dedup testing
|
2022-11-17 16:13:43 +01:00 |
Claudio Atzori
|
3dbc637d3e
|
code formatting
|
2022-11-17 09:55:41 +01:00 |
Claudio Atzori
|
ddff0e8999
|
merging duplicates using IdentifierComparator
|
2022-11-11 16:10:25 +01:00 |
Claudio Atzori
|
5af5a8ae42
|
added IdentifierComparator
|
2022-11-09 14:20:59 +01:00 |
Claudio Atzori
|
7c3390ac10
|
Merge branch 'beta' into eoscifguidelines-from-mdstores
|
2022-11-07 12:18:40 +01:00 |
dimitrispie
|
992fc5b628
|
Added McMaster University Institution
|
2022-11-03 11:02:18 +02:00 |
dimitrispie
|
7fda05e380
|
Added Autonomous University of Barcelona
|
2022-11-01 13:59:40 +02:00 |
Claudio Atzori
|
22873c9172
|
Merge pull request 'Added fields: totalcost, fundedamount, currency, in project table' (#257) from antonis.lempesis/dnet-hadoop:beta into beta
Reviewed-on: D-Net/dnet-hadoop#257
|
2022-10-31 13:49:27 +01:00 |
dimitrispie
|
7861c472e0
|
Hive memory parameters
|
2022-10-28 19:00:32 +03:00 |
dimitrispie
|
5df9c63963
|
Added fields: totalcost, fundedamount, currency, in project table
|
2022-10-27 16:44:26 +03:00 |
Sandro La Bruzzo
|
2b9a20a4a3
|
Changed the way Scholexplorer filter the relationships, I found that filter all relation coming from openCitation is wrong, because we loose a lot of relation than intersect OpenCitation, but they don't come only from there
|
2022-10-24 12:53:47 +02:00 |
Alessia Bardi
|
208ed32315
|
fixed xpath for semantic relation
|
2022-10-23 18:18:13 +02:00 |
Alessia Bardi
|
ee759ac92d
|
file format after mvn compile
|
2022-10-23 18:09:47 +02:00 |
Alessia Bardi
|
31a10f000b
|
Map the field oaf:eoscifguidelines from mdstores. Currently we can find it in ROHub metadata
|
2022-10-23 18:05:37 +02:00 |
Claudio Atzori
|
ec39b84898
|
Merge branch 'beta' of https://code-repo.d4science.org/D-Net/dnet-hadoop into beta
|
2022-10-19 15:21:02 +02:00 |
Claudio Atzori
|
bca4a61710
|
suppressing hyper verbose spark logs during unit test execution
|
2022-10-19 15:20:58 +02:00 |
Sandro La Bruzzo
|
72f0d88d6c
|
formatted code
|
2022-10-19 14:18:42 +02:00 |
Claudio Atzori
|
9b449110c6
|
Merge branch 'beta' of https://code-repo.d4science.org/D-Net/dnet-hadoop into beta
|
2022-10-14 15:48:04 +02:00 |
Claudio Atzori
|
ae7cd0735a
|
[graph2hive] more partitions
|
2022-10-14 15:47:58 +02:00 |
Sandro La Bruzzo
|
135cf81151
|
Merge remote-tracking branch 'origin/beta' into beta
|
2022-10-13 11:47:25 +02:00 |
Sandro La Bruzzo
|
a1f94530a3
|
added documentation
|
2022-10-13 11:47:11 +02:00 |
Claudio Atzori
|
b47aaf4dd1
|
[cleaning] subjects declared as belonging to specific vocabularies whose values are not found in the vocab are set to type keyword
|
2022-10-13 11:23:43 +02:00 |
Claudio Atzori
|
6163ecbf63
|
[cleaning] renamed parameters in wf action
|
2022-10-11 11:20:03 +02:00 |
Claudio Atzori
|
b301e9fdff
|
[cleaning] renamed action name/description
|
2022-10-11 11:08:52 +02:00 |
Claudio Atzori
|
ece40adc09
|
[cleaning] fixing NPE in the country cleaning phase
|
2022-10-11 10:10:20 +02:00 |
Claudio Atzori
|
d51275a965
|
Merge branch 'beta' of https://code-repo.d4science.org/D-Net/dnet-hadoop into beta
|
2022-10-07 09:52:49 +02:00 |
Claudio Atzori
|
8d97949316
|
[cleaning] fixed loop in wf nodes
|
2022-10-07 09:52:45 +02:00 |
Miriam Baglioni
|
a653e1b3ea
|
[Enrichment - result to community through organization] reimplementation of the data preparation step using spark
|
2022-10-04 15:01:28 +02:00 |
Miriam Baglioni
|
4d8339614b
|
Revert "[BipFinder] Fixed issue for wrong escaped char in doi"
This reverts commit 188f25eefa .
|
2022-10-04 14:29:47 +02:00 |
Miriam Baglioni
|
7324853a17
|
Revert "[BipFinder] refactoring"
This reverts commit 28dc317350 .
|
2022-10-04 14:29:39 +02:00 |
Miriam Baglioni
|
28dc317350
|
[BipFinder] refactoring
|
2022-10-04 09:47:27 +02:00 |
Miriam Baglioni
|
188f25eefa
|
[BipFinder] Fixed issue for wrong escaped char in doi
|
2022-10-03 12:42:52 +02:00 |
Claudio Atzori
|
89f7007080
|
Merge pull request '[stats wf] misc changes' (#254) from antonis.lempesis/dnet-hadoop:beta into beta
Reviewed-on: D-Net/dnet-hadoop#254
|
2022-10-03 10:32:05 +02:00 |
dimitrispie
|
2c0c3f1806
|
Cast amount to float for table result_apcs
|
2022-09-28 19:33:24 +03:00 |
Alessia Bardi
|
49360770d7
|
map w3id as instance url
|
2022-09-28 14:16:39 +02:00 |
dimitrispie
|
bdc46e3eaa
|
Remove denormalization of results to fix downloads numbers in monitor
|
2022-09-28 14:59:08 +03:00 |
dimitrispie
|
2ebb1459a9
|
Fixed type in no_downloads
|
2022-09-28 14:36:57 +03:00 |
Miriam Baglioni
|
b5b5a4c192
|
[CleanCountry] fixed issue
|
2022-09-28 12:42:51 +02:00 |
Miriam Baglioni
|
f1d7d45cf7
|
[BulkTag] fixed issue
|
2022-09-28 12:01:43 +02:00 |
Miriam Baglioni
|
3ec044600d
|
[BulkTag] fixed conflicts
|
2022-09-28 11:58:28 +02:00 |
Miriam Baglioni
|
1cb79719a7
|
[BulkTag] fixed issues
|
2022-09-28 11:44:55 +02:00 |
Claudio Atzori
|
f3f7604e6c
|
trying to fix a test that fails only on Jenkins
|
2022-09-27 15:21:37 +02:00 |
Claudio Atzori
|
3f90d159e3
|
code formatting
|
2022-09-27 15:08:00 +02:00 |
Claudio Atzori
|
0b3e44e521
|
Merge branch 'beta' into relation-from-odf
|
2022-09-27 14:57:01 +02:00 |
Claudio Atzori
|
57dbeb08d2
|
code formatting
|
2022-09-27 14:55:10 +02:00 |
Claudio Atzori
|
b60985cf68
|
Merge branch 'beta' into horizontalConstraints
|
2022-09-27 14:39:31 +02:00 |
Claudio Atzori
|
3b60642ef9
|
Merge pull request 'Synchronize indicators in stats-db with monitor-db' (#249) from antonis.lempesis/dnet-hadoop:beta into beta
Reviewed-on: D-Net/dnet-hadoop#249
|
2022-09-27 14:37:33 +02:00 |
Claudio Atzori
|
25e9d92aad
|
Merge branch 'beta' into clean_country
|
2022-09-27 14:27:49 +02:00 |
Alessia Bardi
|
fd63e9bfac
|
Mapping all relationships supported in ModelConstants and ModelSupport
|
2022-09-26 11:24:13 +02:00 |
Miriam Baglioni
|
ca216a92ad
|
[BulkTagging] changed the query to the IS to insert values for FOS and SDG as subject in the configuration used for the tagging
|
2022-09-23 17:06:07 +02:00 |
Miriam Baglioni
|
3e6b0f58bb
|
[BulkTagging] changed the query to the IS to get also the information for the advancedConstraint from the profile
|
2022-09-23 16:47:19 +02:00 |
Miriam Baglioni
|
4a3e119b73
|
mergin with branch beta
|
2022-09-23 16:16:06 +02:00 |
Miriam Baglioni
|
f0e303abf9
|
Merge branch 'beta' of https://code-repo.d4science.org/D-Net/dnet-hadoop into beta
|
2022-09-23 16:15:32 +02:00 |
Miriam Baglioni
|
55da4d8715
|
[BulkTagging] modifying code to represent constraints horizontally on all the results. Added subject to the set of field used to express the constraint. Modified resorces to test the new approach. Modified test calss
|
2022-09-23 16:02:19 +02:00 |
Alessia Bardi
|
c5eb722170
|
relationships from relatedIdentifier whose target id type is one of the pid type with an authority
|
2022-09-23 15:47:05 +02:00 |
Claudio Atzori
|
c86cc53520
|
suppressing hyper verbose spark logs during unit test execution
|
2022-09-23 15:20:40 +02:00 |
Alessia Bardi
|
ba33ff71fd
|
refactoring for the generation of relationships from related identifier of type 'OPENAIRE'
|
2022-09-23 15:17:13 +02:00 |
Alessia Bardi
|
982bcc1e35
|
test wrid pid and record identifier
|
2022-09-23 12:06:06 +02:00 |
Miriam Baglioni
|
960cb861a0
|
refactoring
|
2022-09-23 11:14:04 +02:00 |
Claudio Atzori
|
c42850328e
|
fixed semantic (subreltype) for ServiceOrganization relations
|
2022-09-22 16:23:25 +02:00 |
Miriam Baglioni
|
33bb79459e
|
Merge branch 'beta' of https://code-repo.d4science.org/D-Net/dnet-hadoop into beta
|
2022-09-22 15:55:17 +02:00 |
dimitrispie
|
dcd85f8cd7
|
- Synchronize indicators in stats-db with monitor-db
- added new openorg id for Nanyang Technological University
- changed openorg id for University of Helsinki #8088 ticket
|
2022-09-22 13:33:07 +03:00 |
Claudio Atzori
|
e45ec15221
|
Merge branch 'beta' into clean_country
|
2022-09-19 11:34:02 +02:00 |
Claudio Atzori
|
26e1badded
|
added instance.url syntactical validation, avoid creating multiple duplicated URLs
|
2022-09-19 11:19:10 +02:00 |
Miriam Baglioni
|
5240ac3d7b
|
[EOSC Tag] remove addition of eosc context for result with eosc if guidelines set
|
2022-09-19 11:02:18 +02:00 |
Claudio Atzori
|
192215a18e
|
merged from branch discard-non-wellformed
|
2022-09-19 10:17:10 +02:00 |
Claudio Atzori
|
e370e940d8
|
[aggregator graph] save invalid records aside for further inspection
|
2022-09-16 14:06:28 +02:00 |
Claudio Atzori
|
465e941214
|
Merge pull request '[stats wf] Changes to indicators tables' (#244) from antonis.lempesis/dnet-hadoop:beta into beta
Reviewed-on: D-Net/dnet-hadoop#244
|
2022-09-16 10:13:58 +02:00 |
Claudio Atzori
|
1e42d984e1
|
[aggregator graph] save invalid records aside for further inspection
|
2022-09-15 10:49:42 +02:00 |
Alessia Bardi
|
9e7ec4198f
|
fixed test
|
2022-09-14 18:08:56 +02:00 |
Claudio Atzori
|
c48f6e9c57
|
[aggregator graph] save invalid records aside for further inspection
|
2022-09-14 17:11:26 +02:00 |
dimitrispie
|
3bf3127251
|
Changes to monitor and indicator scripts
|
2022-09-14 16:36:19 +03:00 |
Claudio Atzori
|
a0919ed495
|
[aggregator graph] save invalid records aside for further inspection
|
2022-09-14 13:27:39 +02:00 |
Alessia Bardi
|
b99a011345
|
return empty Oaf list if record cannot be parsed
|
2022-09-13 11:51:55 +02:00 |
Alessia Bardi
|
27af5122d2
|
logs for non well formed XML files
|
2022-09-12 14:25:23 +02:00 |
Claudio Atzori
|
ff6f789b6d
|
code formatting
|
2022-09-09 15:16:31 +02:00 |
Claudio Atzori
|
b5d6966c01
|
Merge branch 'beta' into clean_country
|
2022-09-09 12:20:19 +02:00 |
Claudio Atzori
|
b5f7bd30be
|
Merge branch 'beta' into clean_subjects
|
2022-09-09 12:20:04 +02:00 |
Alessia Bardi
|
f14107ad77
|
Merge branch 'handle_as_instance_urls' of https://code-repo.d4science.org/D-Net/dnet-hadoop into handle_as_instance_urls
|
2022-09-09 12:17:19 +02:00 |
Alessia Bardi
|
a539c6ccaf
|
https for handle URLs
|
2022-09-09 12:16:28 +02:00 |
dimitrispie
|
71b069ca90
|
Changes to indicator and monitor scripts
|
2022-09-09 13:15:58 +03:00 |
Claudio Atzori
|
1203378441
|
Merge branch 'beta' into clean_subjects
|
2022-09-09 10:38:47 +02:00 |
Claudio Atzori
|
14dc909a14
|
Merge branch 'beta' into clean_country
|
2022-09-09 10:38:17 +02:00 |
Claudio Atzori
|
853c996fa2
|
Merge branch 'beta' into handle_as_instance_urls
|
2022-09-09 09:47:16 +02:00 |
Claudio Atzori
|
a431e01383
|
Merge pull request 'orcid_multipleworks_download' (#242) from enrico.ottonello/dnet-hadoop:orcid_multipleworks_download into beta
Reviewed-on: D-Net/dnet-hadoop#242
|
2022-09-09 08:45:02 +02:00 |
Alessia Bardi
|
9ef063d502
|
#7861#note-8 instance url from handle
|
2022-09-07 17:29:54 +03:00 |
Alessia Bardi
|
5c45d52af3
|
testing for RiuNet
|
2022-09-07 15:40:57 +03:00 |
dimitrispie
|
2b5f8c9c9a
|
comment out duplicate table creation
|
2022-09-06 12:27:53 +03:00 |
Alessia Bardi
|
a11eb38065
|
testing for RO-Hub
|
2022-09-02 16:07:36 +02:00 |
Enrico Ottonello
|
bfdf2dc390
|
Merge branch 'beta' of https://code-repo.d4science.org/D-Net/dnet-hadoop into orcid_multipleworks_download
|
2022-08-25 12:07:54 +02:00 |
Enrico Ottonello
|
da1cf561e6
|
alignment with beta
|
2022-08-25 11:57:20 +02:00 |
Enrico Ottonello
|
27445ccdaa
|
cleaned log
|
2022-08-25 11:56:14 +02:00 |
Claudio Atzori
|
b7c387c21f
|
cleaning of subjects: avoid duplicated subjects, prioritise collected vs inferred or other sources
|
2022-08-12 15:09:16 +02:00 |
Claudio Atzori
|
adb526b0e1
|
Merge branch 'beta' into clean_subjects
|
2022-08-12 10:51:17 +02:00 |
Claudio Atzori
|
cb7c07c54e
|
[scholix] added step to create tar archive
|
2022-08-11 11:25:24 +02:00 |
Claudio Atzori
|
2aa16d0432
|
[scholix] fixed OpenCitation dump procedure
|
2022-08-10 17:39:29 +02:00 |
Miriam Baglioni
|
7dbdd4a0fe
|
[Clean Country]changes related to D-Net/dnet-hadoop#241 (comment)
|
2022-08-10 15:13:10 +02:00 |
Claudio Atzori
|
51ad93e545
|
[scholix] fixed OpenCitation dump procedure
|
2022-08-10 11:57:56 +02:00 |
Miriam Baglioni
|
62d2138806
|
[Clean Context] changed a bit the logic. Added the check not to have result hosted by a datasource of type institutional repository from NL. Added also the check that the country should have been included in the result via propagation for it to be removed
|
2022-08-08 14:10:47 +02:00 |
Claudio Atzori
|
3418ce50ac
|
cleaning of subjects: perform the cleaning when the given value is equivalent to one of the terms in the vocabulary
|
2022-08-08 12:48:47 +02:00 |
Claudio Atzori
|
a78028dabc
|
Merge branch 'beta' into clean_subjects
|
2022-08-08 12:34:33 +02:00 |
Miriam Baglioni
|
390013a4b2
|
mergin with branch beta
|
2022-08-08 12:30:31 +02:00 |
Claudio Atzori
|
3937ff04de
|
Merge branch 'beta' into tagEosc
|
2022-08-08 09:57:23 +02:00 |
Claudio Atzori
|
a4815f6bec
|
Merge branch 'beta' into clean_subjects
|
2022-08-05 16:57:03 +02:00 |
Claudio Atzori
|
29c4cde42e
|
Merge branch 'clean_subjects' of https://code-repo.d4science.org/D-Net/dnet-hadoop into clean_subjects
|
2022-08-05 16:56:37 +02:00 |
Claudio Atzori
|
4eaa063b1f
|
cleaning of subjects
|
2022-08-05 16:56:09 +02:00 |
Claudio Atzori
|
84598c7535
|
Merge pull request 'restored some collab indicators' (#240) from antonis.lempesis/dnet-hadoop:beta into beta
Reviewed-on: D-Net/dnet-hadoop#240
|
2022-08-05 15:50:39 +02:00 |
Antonis Lempesis
|
fcef5294e2
|
restored some collab indicators
|
2022-08-05 13:45:01 +03:00 |
Claudio Atzori
|
844f6eb465
|
Merge branch 'beta' into clean_subjects
|
2022-08-05 12:39:05 +02:00 |
Claudio Atzori
|
32cee1f619
|
WIP: cleaning of subjects
|
2022-08-05 12:32:08 +02:00 |
Claudio Atzori
|
c1f2ffc53d
|
Merge pull request 'commenting out the collab indicators because they still fail' (#237) from antonis.lempesis/dnet-hadoop:beta into beta
Reviewed-on: D-Net/dnet-hadoop#237
|
2022-08-05 11:57:36 +02:00 |
Antonis Lempesis
|
227e10f4b3
|
commenting out the collab indicators because they still fail
|
2022-08-05 12:54:36 +03:00 |
Claudio Atzori
|
6c0fd9284b
|
merge from beta
|
2022-08-05 10:42:53 +02:00 |
Claudio Atzori
|
b78889a0ce
|
WIP: cleaning of subjects
|
2022-08-05 09:11:37 +02:00 |
Miriam Baglioni
|
a7a18d7630
|
[Graph Dump] removed code for the dump from the project. Fixed issues in tests when possible
|
2022-08-04 17:40:40 +02:00 |
Claudio Atzori
|
499826ead1
|
serialising field eoscifguidelines field in the Solr XML records
|
2022-08-04 12:40:48 +02:00 |
Claudio Atzori
|
27a91841e7
|
WIP: cleaning of subjects
|
2022-08-04 11:39:39 +02:00 |
Antonis Lempesis
|
b09d7ddc74
|
fixed the datasourceOrganization relations
|
2022-08-03 12:26:50 +02:00 |
Claudio Atzori
|
e62018e95d
|
[aggregator graph] added more assertions in test
|
2022-08-03 12:26:05 +02:00 |
Claudio Atzori
|
efd96e7e66
|
Merge pull request 'fixed the datasourceOrganization relations' (#233) from antonis.lempesis/dnet-hadoop:beta into beta
Reviewed-on: D-Net/dnet-hadoop#233
|
2022-08-03 12:25:05 +02:00 |
Antonis Lempesis
|
8b0407d8ec
|
fixed the datasourceOrganization relations
|
2022-08-03 12:26:59 +03:00 |
Claudio Atzori
|
eb53b52f7c
|
code formatting
|
2022-08-02 13:24:47 +02:00 |
Claudio Atzori
|
27681cf6bf
|
Merge pull request '[stats wf] latest version of indicators + added FOS classification' (#232) from antonis.lempesis/dnet-hadoop:beta into beta
Reviewed-on: D-Net/dnet-hadoop#232
|
2022-08-02 12:57:15 +02:00 |
Antonis Lempesis
|
1778d40c40
|
latest version of indicators
|
2022-08-02 13:39:34 +03:00 |
Claudio Atzori
|
209c7e9dab
|
[datacite] avoid UnsupportedOperationException
|
2022-08-01 09:05:35 +02:00 |
Enrico Ottonello
|
64311b8be4
|
removed unuseful accumulator
|
2022-07-31 01:03:29 +02:00 |
Antonis Lempesis
|
6fc9ef53f6
|
addded command line params to allow hive actions to run
|
2022-07-29 16:36:20 +03:00 |
Antonis Lempesis
|
9886fe87ec
|
- Added FOS classification
- Added extra orgs in monitor
- Fixed result-project and organization-project tables
|
2022-07-29 16:34:50 +03:00 |
Claudio Atzori
|
92e48f12f7
|
[metadata collection] updated collector plugin name
|
2022-07-29 13:54:00 +02:00 |
Claudio Atzori
|
f62c4e05cd
|
code formatting
|
2022-07-29 11:56:01 +02:00 |
Claudio Atzori
|
0727f0ef48
|
[EOSC tag] avoid NPEs
|
2022-07-29 11:55:34 +02:00 |
Miriam Baglioni
|
3329b6ce6b
|
[EOSC TAG] added fix for NPE on subjects
|
2022-07-29 10:54:20 +02:00 |
Claudio Atzori
|
1dd1e4fe3a
|
extended test for mapping project_organization relations
|
2022-07-28 11:27:08 +02:00 |
Claudio Atzori
|
60e4fbd78b
|
Merge branch 'beta' into project_organization_contribution
|
2022-07-28 10:15:43 +02:00 |
Claudio Atzori
|
ed98a6d9d0
|
[Datacite mapping] include the older datacite prefixed OpenAIRE id among the originalId[]
|
2022-07-28 10:15:14 +02:00 |
Claudio Atzori
|
09ccc7b472
|
Merge branch 'beta' into project_organization_contribution
|
2022-07-28 09:49:59 +02:00 |
Sandro La Bruzzo
|
67525076ec
|
fixed test, now it compiles after commit a6977197b3
|
2022-07-26 15:35:17 +02:00 |
Claudio Atzori
|
26104826c4
|
Merge branch 'beta' of https://code-repo.d4science.org/D-Net/dnet-hadoop into beta
|
2022-07-26 14:34:29 +02:00 |
Claudio Atzori
|
d43663d30f
|
adapted RorActionSet test, it should not create parent/child rels
|
2022-07-25 17:54:10 +02:00 |
Miriam Baglioni
|
35bcd9422d
|
[EOSC Context Tagging] removed not needed specification in path
|
2022-07-25 15:45:22 +02:00 |
Miriam Baglioni
|
1c82acb168
|
[EOSC Context Tagging] refactoring: moved EOSC IF tagging in package eosc under bulkTag
|
2022-07-25 14:26:39 +02:00 |
Miriam Baglioni
|
68cb637832
|
merge with branch beta
|
2022-07-25 14:24:25 +02:00 |
Miriam Baglioni
|
0172bab251
|
[EOSC Context Tagging] refactoring
|
2022-07-25 14:16:45 +02:00 |
Claudio Atzori
|
612b7a5530
|
Merge branch 'beta' into tagEosc
|
2022-07-25 14:12:59 +02:00 |
Claudio Atzori
|
c3ede1b379
|
Merge branch 'beta' into pubmed_update
|
2022-07-25 14:10:22 +02:00 |
Miriam Baglioni
|
144c103b67
|
[EOSC Context Tagging] add check to avoid the insertion of the context if already present
|
2022-07-25 13:52:45 +02:00 |
Enrico Ottonello
|
657b0208a2
|
multiple works download (<=100) for single request
|
2022-07-25 12:37:39 +02:00 |
Miriam Baglioni
|
d091866e48
|
[EOSC Context Tagging] refactoring
|
2022-07-25 11:12:22 +02:00 |
Miriam Baglioni
|
5968ec018d
|
[Clean Country] modified workflow and added param file
|
2022-07-22 16:48:38 +02:00 |
Miriam Baglioni
|
a12d28c644
|
[Clean Country] added logic not to remove country from result if it exist a hosting datasource with that country. Moreover the country will be removed only if added with propagation
|
2022-07-22 16:23:12 +02:00 |
Miriam Baglioni
|
2c933f1158
|
mergin with branch beta
|
2022-07-22 14:57:41 +02:00 |
Miriam Baglioni
|
06a95daf60
|
[EOSC context TAG] refactoring after compilation
|
2022-07-22 14:57:06 +02:00 |
Miriam Baglioni
|
ffb0ce3fb9
|
mergin with branch beta
|
2022-07-22 14:55:55 +02:00 |
Miriam Baglioni
|
627332526b
|
[EOSC context TAG] workflow start from reset_outputpath action
|
2022-07-22 14:55:11 +02:00 |
Miriam Baglioni
|
7a1c1b6f53
|
[EOSC context TAG] Add test class and resourcesK
|
2022-07-22 14:36:02 +02:00 |
Sandro La Bruzzo
|
ddc414b258
|
fixed wrong json param
|
2022-07-22 09:43:15 +02:00 |
Miriam Baglioni
|
317a4a56ef
|
[EOSC context TAG] first implementation of the logic to tag results imported from datasources registered in the EOSC
|
2022-07-21 17:37:48 +02:00 |
Miriam Baglioni
|
3be036f290
|
[EOSC TAG] refactoring after compilation
|
2022-07-21 14:45:43 +02:00 |
Miriam Baglioni
|
e61b8e6b03
|
mergin with branch beta
|
2022-07-21 14:43:23 +02:00 |
Miriam Baglioni
|
56d09e6348
|
[EOSC TAG] before adding the tag added a step to verify the same tag is not already present
|
2022-07-21 14:36:48 +02:00 |
Miriam Baglioni
|
5143a80232
|
[EOSC TAG] modification of test class to align with new element
|
2022-07-21 11:56:51 +02:00 |
Sandro La Bruzzo
|
5f651f2316
|
changed filter relation on SubRelType
|
2022-07-21 10:11:48 +02:00 |
Miriam Baglioni
|
438abdf96f
|
[EOSC TAG] adding eosc interoperability guidelines in the specific element in the result. Removed from subjects. Removed also the deletion of EOSC Jupyter Notebook from subject since now the criteria are searchd for in a different place
|
2022-07-20 18:07:54 +02:00 |
Miriam Baglioni
|
65cc736e2f
|
[Clean Country] first implementation to remove country NL from results collected from NARCIS when doi starts with mendely prefix
|
2022-07-20 17:05:56 +02:00 |
Sandro La Bruzzo
|
5b76321d9c
|
implemented oozie workflow to generate scholix dump filtering relclass semantic
|
2022-07-20 16:34:32 +02:00 |
Claudio Atzori
|
1138b2ac8e
|
code formatting
|
2022-07-19 14:15:49 +02:00 |
Sandro La Bruzzo
|
00168303db
|
Added unit test to verify the generation in the OriginalID the old openaire Identifier generated by OAI
|
2022-07-14 10:19:59 +02:00 |
Sandro La Bruzzo
|
0a4f4d98fa
|
added PMCId to PmArticle
|
2022-07-13 15:27:17 +02:00 |
Claudio Atzori
|
0c1cfee396
|
mapping oaf:fulltext elements in the result.fulltext field
|
2022-07-11 17:34:59 +02:00 |
Miriam Baglioni
|
fae681fea1
|
[Country Propagation] add check to avoid NPE on datasource.getDatasourceType().getClassis()
|
2022-07-03 17:39:58 +02:00 |
Miriam Baglioni
|
c09fcdb40b
|
Merge branch 'beta' of https://code-repo.d4science.org/D-Net/dnet-hadoop into beta
|
2022-07-01 12:38:03 +02:00 |
Claudio Atzori
|
0cb1c70788
|
code formatting
|
2022-07-01 10:44:08 +02:00 |
Claudio Atzori
|
4ec13e2b66
|
Merge branch 'master' into dump_new_funded_products
|
2022-07-01 10:30:28 +02:00 |
Claudio Atzori
|
072f192853
|
include the class information in the measure XML serialization
|
2022-07-01 09:54:56 +02:00 |
Claudio Atzori
|
a88103bcf9
|
[action manager] added more testing
|
2022-07-01 09:06:59 +02:00 |
Claudio Atzori
|
7da24c1dec
|
added more logging
|
2022-06-28 13:47:49 +02:00 |
Miriam Baglioni
|
ee1f1eeca2
|
Merge branch 'beta' of https://code-repo.d4science.org/D-Net/dnet-hadoop into beta
|
2022-06-28 11:06:32 +02:00 |
Miriam Baglioni
|
71744a1f52
|
[DUMP DELTA PROJECTS] refactoring
|
2022-06-27 18:07:58 +02:00 |
Miriam Baglioni
|
1d1fe3b151
|
[DUMP DELTA PROJECTS] refactoring
|
2022-06-27 18:04:59 +02:00 |
Claudio Atzori
|
a8773af0cb
|
Merge branch 'beta' into project_organization_contribution
|
2022-06-27 09:37:40 +02:00 |
Claudio Atzori
|
4829b96bb5
|
Merge branch 'beta' into author_name_particles
|
2022-06-27 09:37:03 +02:00 |
Claudio Atzori
|
5130eac247
|
mapping by participant project contribution
|
2022-06-24 17:16:42 +02:00 |
Claudio Atzori
|
929b145130
|
code formatting
|
2022-06-21 23:07:06 +02:00 |
Miriam Baglioni
|
edddfc6c63
|
[DUMP DELTA PROJECTS] adding test and resource
|
2022-06-21 18:28:53 +02:00 |
Miriam Baglioni
|
f561f13dd9
|
[Funder Products Dump] fixed names of parameters in workflow
|
2022-06-21 18:18:17 +02:00 |
Miriam Baglioni
|
ff74e73369
|
[DUMP NEW FUNDED PRODUCTS] change in resources
|
2022-06-21 18:02:51 +02:00 |
Miriam Baglioni
|
b98f904d48
|
[Funder Products Dump] new way to avoid using hive
|
2022-06-21 17:52:27 +02:00 |
Miriam Baglioni
|
7423577a08
|
[Graph DUMP] add code to produce the delta of new projects with respect to the previous delta/dump
|
2022-06-21 14:51:38 +02:00 |
Claudio Atzori
|
b295a40d9c
|
restored use of name_particles when parsing author names
|
2022-06-16 12:20:43 +02:00 |
Claudio Atzori
|
c7b09c6225
|
Merge branch 'beta' into 7096-fileGZip-collector-plugin
|
2022-06-16 09:28:50 +02:00 |
Claudio Atzori
|
e03c0c7794
|
Merge branch 'beta' into oaf_relation_mapping
|
2022-06-16 09:27:01 +02:00 |
Claudio Atzori
|
06b5533d4c
|
Merge branch 'beta' into 7096-fileGZip-collector-plugin
|
2022-06-16 09:22:16 +02:00 |
Claudio Atzori
|
4c8e820ff0
|
mapping relationship from trasformed records based on oaf:relation
|
2022-06-14 08:49:02 +02:00 |
Alessia Bardi
|
88d531dc91
|
exclude FAIRsharing records from Datacite
|
2022-06-13 16:17:17 +02:00 |
Claudio Atzori
|
116902c028
|
mapping relationship from trasformed records based on oaf:relation
|
2022-06-13 14:31:48 +02:00 |
Claudio Atzori
|
b8cda65487
|
code formatting
|
2022-06-13 09:20:03 +02:00 |
Michele Artini
|
634869ce95
|
deleted hierarchical rels from ror action set
|
2022-06-13 09:12:21 +02:00 |
Alessia Bardi
|
922c6d66ef
|
Merge branch 'beta' of https://code-repo.d4science.org/D-Net/dnet-hadoop into beta
|
2022-06-10 17:29:15 +02:00 |
Alessia Bardi
|
68bd58d6a4
|
tests for ROHub
|
2022-06-10 17:29:11 +02:00 |
Miriam Baglioni
|
b229c6e7af
|
Merge pull request 'beta' (#218) from antonis.lempesis/dnet-hadoop:beta into beta
Reviewed-on: D-Net/dnet-hadoop#218
|
2022-06-10 11:03:48 +02:00 |
Antonis Lempesis
|
ab18c9daa9
|
Merge branch 'beta' of https://code-repo.d4science.org/antonis.lempesis/dnet-hadoop into beta
|
2022-06-09 15:48:21 +03:00 |
Antonis Lempesis
|
574492c659
|
removed double result_apc table creation from monitor
|
2022-06-09 15:48:13 +03:00 |
Michele Artini
|
b94a791bc5
|
unit tests to transform cnr explora
|
2022-06-09 12:25:34 +02:00 |
Miriam Baglioni
|
4b6913787b
|
[DOI-BOOST] added one method in test of crossref mapping to aof and one resource. Related to ticket 7807
|
2022-06-08 14:55:19 +02:00 |
Antonis Lempesis
|
db088cc69c
|
fixed *_organization tables
|
2022-06-07 04:04:28 +03:00 |
Miriam Baglioni
|
31d4557e8d
|
Merge branch 'master' of https://code-repo.d4science.org/D-Net/dnet-hadoop
|
2022-06-06 11:52:29 +02:00 |
Claudio Atzori
|
5c2949a864
|
Merge pull request '[stats wf] added open citations & more orgs in monitor, removed collab indicator' (#213) from antonis.lempesis/dnet-hadoop:beta into beta
Reviewed-on: D-Net/dnet-hadoop#213
|
2022-05-20 11:38:43 +02:00 |
Miriam Baglioni
|
5e0b8f9b5f
|
[CountryPropagation] refactoring
|
2022-05-20 09:15:53 +02:00 |
Miriam Baglioni
|
c298c148cb
|
[CountryPropagation] fix NPE issue
|
2022-05-20 09:11:46 +02:00 |
Miriam Baglioni
|
eaf9385ae5
|
Merge branch 'beta' of https://code-repo.d4science.org/D-Net/dnet-hadoop into beta
|
2022-05-17 15:09:37 +02:00 |
Miriam Baglioni
|
f5207885e3
|
[EOSCTag] changed code to remove EOSC Jupyter Notebook and modified test to exclude galaxy + software from the tagging for Galaxy
|
2022-05-17 15:09:22 +02:00 |
Claudio Atzori
|
d098ad0d93
|
[hb patch] updated map
|
2022-05-16 15:54:04 +02:00 |
Claudio Atzori
|
1dda11e031
|
[hb patch] updated map
|
2022-05-16 15:53:27 +02:00 |
Claudio Atzori
|
8dd5517548
|
code formatting
|
2022-05-16 14:35:24 +02:00 |
Claudio Atzori
|
52cb086506
|
[graph grouping] drop relation target path before copying from source
|
2022-05-16 12:08:36 +02:00 |
Claudio Atzori
|
6442763f97
|
Merge branch 'beta' of https://code-repo.d4science.org/D-Net/dnet-hadoop into beta
|
2022-05-16 12:07:45 +02:00 |
Claudio Atzori
|
997c50078e
|
[graph grouping] drop relation target path before copying from source
|
2022-05-16 12:07:40 +02:00 |
Sandro La Bruzzo
|
c1971d52c4
|
Merge branch 'beta' of code-repo.d4science.org:D-Net/dnet-hadoop into beta
|
2022-05-16 10:30:35 +02:00 |
Sandro La Bruzzo
|
4c50f35c8b
|
update publication Date format
|
2022-05-16 10:29:36 +02:00 |
Michele Artini
|
46c07e0724
|
deleted hierarchical rels from ror action set
|
2022-05-16 09:39:54 +02:00 |
Claudio Atzori
|
6031acb2e3
|
[openorgs] fixed parent/child query, using the correct semantic labels
|
2022-05-16 09:20:48 +02:00 |
Claudio Atzori
|
0dc33ea391
|
[openorgs] fixed parent/child query, using the correct semantic labels
|
2022-05-16 09:20:30 +02:00 |
Antonis Lempesis
|
3fc9efeab6
|
fixed typo, addded open citations and apcs in monitor
|
2022-05-13 14:28:13 +03:00 |
Miriam Baglioni
|
e4eac1d20b
|
[EOSC TAG] added code to remove EOSC Jupyter Notebook from subjects and put EOSC as classid in the qualifier
|
2022-05-13 11:01:33 +02:00 |
Sandro La Bruzzo
|
22f65680b9
|
Merge branch 'beta' of code-repo.d4science.org:D-Net/dnet-hadoop into beta
|
2022-05-11 15:30:12 +02:00 |
Sandro La Bruzzo
|
ca8d26bcb4
|
added better filter for openCitations
|
2022-05-11 15:29:57 +02:00 |
Claudio Atzori
|
5d3b4a9c25
|
[graph merge beta] merge datasource originalid, collectedfrom, and pid lists
|
2022-05-11 14:13:06 +02:00 |
Antonis Lempesis
|
23334479bb
|
removed yet another collab, added more orgs in monitor
|
2022-05-11 13:05:52 +03:00 |
Claudio Atzori
|
2a8e0fb72f
|
[openorgs] mapping parent/child relations without massaging the semantic labels
|
2022-05-10 08:45:53 +02:00 |
Claudio Atzori
|
77bc9863e9
|
[openorgs] mapping parent/child relations without massaging the semantic labels
|
2022-05-09 16:06:04 +02:00 |
Claudio Atzori
|
378020e30a
|
[eosc_services] unit test adaptation
|
2022-05-09 16:05:06 +02:00 |
Miriam Baglioni
|
89657a0b78
|
[UsageCount] refactoring
|
2022-05-09 14:43:27 +02:00 |
Miriam Baglioni
|
a056f59c6e
|
[UsageCount] make it as an action set as it should be, plus changed the test to make them work as well now
|
2022-05-09 12:51:35 +02:00 |
Antonis Lempesis
|
61b4c19e65
|
restored indi_result_org_country_collab, removed indi_result_org_collab
|
2022-05-06 12:52:10 +03:00 |
Antonis Lempesis
|
cfbbcaf7c4
|
commented out indi_result_org_country_collab
|
2022-05-06 12:49:36 +03:00 |
Claudio Atzori
|
658450d9a3
|
Merge branch 'beta' of https://code-repo.d4science.org/D-Net/dnet-hadoop into beta
|
2022-05-05 11:38:08 +02:00 |
Claudio Atzori
|
846975c886
|
[eosc_services] using the correct 'keyword' subject type, as declared in the dnet:subject_classification_typologies vocabulary
|
2022-05-05 11:37:58 +02:00 |
Miriam Baglioni
|
8a72de4011
|
[EOSCTag] modified workflow to execute all the steps and not only the last one
|
2022-05-04 10:10:56 +02:00 |
Miriam Baglioni
|
bd1108f98b
|
mergin with branch beta
|
2022-05-04 10:06:56 +02:00 |
Miriam Baglioni
|
3aeedd931a
|
[EOSCTag] fixed issue in case description is null. Modified test resources and classes
|
2022-05-04 10:06:38 +02:00 |
Claudio Atzori
|
da611cfbbd
|
[eosc_services] resolved merge conflicts
|
2022-05-03 13:37:15 +02:00 |
Claudio Atzori
|
9e12cb3c92
|
EOSC Services - removed field knowledgegraph; depending on the released schema module
|
2022-05-03 11:55:45 +02:00 |
Miriam Baglioni
|
a21fe310e5
|
[EOSCTag] last test and change in the implementation to search in title and descriptio
|
2022-05-02 17:43:20 +02:00 |
Claudio Atzori
|
2ade69dea6
|
EOSC Services - minor
|
2022-05-02 17:03:31 +02:00 |
Claudio Atzori
|
b6a7ff3a99
|
EOSC Services - removed fields from mapping, testing preparation
|
2022-05-02 15:52:33 +02:00 |
Miriam Baglioni
|
e37177e1ce
|
mergin with branch beta
|
2022-05-02 12:31:50 +02:00 |
Claudio Atzori
|
a8c51f6f16
|
EOSC Services - fixed query and testing preparation
|
2022-05-02 11:09:03 +02:00 |
Claudio Atzori
|
05c1ea92e9
|
EOSC Services - added Service-specific fields in the XML record serialization
|
2022-04-29 15:56:55 +02:00 |
Claudio Atzori
|
f5f532d134
|
EOSC Services - ongoing update
|
2022-04-29 12:25:24 +02:00 |
Antonis Lempesis
|
0353f93d54
|
added new hive opts
|
2022-04-29 12:49:27 +03:00 |
Serafeim Chatzopoulos
|
623f7be26d
|
Fix reading files from HDFS in FileCollector & FileGZipCollector plugins
|
2022-04-28 16:31:11 +03:00 |
Claudio Atzori
|
5ffc24d1ba
|
EOSC Services - ongoing update
|
2022-04-26 16:18:41 +02:00 |
Sandro La Bruzzo
|
78015a5733
|
Merge branch 'beta' of code-repo.d4science.org:D-Net/dnet-hadoop into beta
|
2022-04-26 09:56:34 +02:00 |
Sandro La Bruzzo
|
8c22e5c30a
|
added fix to include date array with only year or year and month
|
2022-04-26 09:56:27 +02:00 |
Claudio Atzori
|
81c4496d32
|
Merge branch 'beta' into 7096-fileGZip-collector-plugin
|
2022-04-26 09:02:15 +02:00 |
Miriam Baglioni
|
e342ec93f0
|
[EOSCTag] prepared resources for test
|
2022-04-22 18:35:37 +02:00 |
Miriam Baglioni
|
88562c0930
|
[EOSC TAG] added test for galaxy for title and description criterias
|
2022-04-22 18:35:03 +02:00 |
Miriam Baglioni
|
dfbd2bcbea
|
[EOSC TAG] added logic in case subject is null
|
2022-04-22 18:34:03 +02:00 |
Miriam Baglioni
|
27c85e901a
|
[EOSCTag] added resources and finalized test for Jupyter Notebook tagging
|
2022-04-22 17:38:10 +02:00 |
Miriam Baglioni
|
87bff36d9e
|
mergin with branch beta
|
2022-04-22 15:52:34 +02:00 |
Miriam Baglioni
|
911ce0780a
|
Merge branch 'cleancontext' of https://code-repo.d4science.org/D-Net/dnet-hadoop into cleancontext
|
2022-04-22 15:41:42 +02:00 |
Miriam Baglioni
|
19d90658fc
|
[Clean Context] added description to parameters
|
2022-04-22 15:41:23 +02:00 |
Claudio Atzori
|
54162f5c4f
|
Merge branch 'beta' into cleancontext
|
2022-04-22 11:49:33 +02:00 |
Miriam Baglioni
|
bbb77052d3
|
[EOSCTag] first test
|
2022-04-22 11:32:57 +02:00 |
Claudio Atzori
|
30105f0722
|
Merge branch 'beta' into 7096-fileGZip-collector-plugin
|
2022-04-22 11:22:21 +02:00 |
Sandro La Bruzzo
|
a82ec3aaaf
|
code formatter
|
2022-04-22 11:08:13 +02:00 |
Sandro La Bruzzo
|
aa12429f50
|
Modified last intersection since we lost many titles.
|
2022-04-22 11:05:08 +02:00 |
Miriam Baglioni
|
7cb7066472
|
[EoscTag] first "rough" implementation
|
2022-04-22 10:44:17 +02:00 |
Sandro La Bruzzo
|
d660895b30
|
fixed wrong mapping type of dataset
|
2022-04-21 20:41:13 +02:00 |
Miriam Baglioni
|
e0915061c2
|
[Clean Context] fixed issue in param name
|
2022-04-21 16:32:40 +02:00 |
Miriam Baglioni
|
6dc68c48e0
|
[EOSCTag] -
|
2022-04-21 16:19:04 +02:00 |
Miriam Baglioni
|
9a961a0092
|
[Clean Context] fixed issue in param name
|
2022-04-21 15:12:24 +02:00 |
Claudio Atzori
|
29150a5d0c
|
code formatting
|
2022-04-21 13:31:56 +02:00 |
Miriam Baglioni
|
5b7d9e741c
|
[Clean Context] added logic to cleaning workflow to accomodate also context cleaning
|
2022-04-21 13:02:14 +02:00 |
Miriam Baglioni
|
ccba1a3db1
|
[Clean Context] added logic to cleaning workflow to accomodate also context cleaning
|
2022-04-21 13:00:06 +02:00 |
Miriam Baglioni
|
20de75ca64
|
[Measures] removed typo
|
2022-04-21 12:14:03 +02:00 |
Miriam Baglioni
|
bebb2a0560
|
Merge branch 'eosc_dimitris' of https://code-repo.d4science.org/D-Net/dnet-hadoop into eosc_dimitris
|
2022-04-21 12:10:19 +02:00 |
Miriam Baglioni
|
b61efd613b
|
[Measures] addressed comments in the PR
|
2022-04-21 12:09:37 +02:00 |
Miriam Baglioni
|
d012d125d7
|
[EOSCTag] -
|
2022-04-21 12:02:09 +02:00 |
Claudio Atzori
|
88acad76f9
|
Merge branch 'beta' into eosc_dimitris
|
2022-04-21 12:00:03 +02:00 |
Claudio Atzori
|
eabb40fccc
|
Merge branch 'beta' into 7096-fileGZip-collector-plugin
|
2022-04-21 11:42:43 +02:00 |
Miriam Baglioni
|
c304657d91
|
[Measures] put the logic in common, no need to change the schema
|
2022-04-21 11:27:26 +02:00 |
Sandro La Bruzzo
|
d580e15442
|
Modified last intersection since we lost many titles.
this is my last resource, after that, I've to change my job
|
2022-04-21 11:06:08 +02:00 |
Miriam Baglioni
|
5295effc96
|
[Measures] fixed issue
|
2022-04-20 16:20:40 +02:00 |
Miriam Baglioni
|
a38f0f5ea7
|
mergin with branch beta
|
2022-04-20 15:44:18 +02:00 |
Miriam Baglioni
|
dbfbe8841a
|
[Clean Context] changed the description in input parameters
|
2022-04-20 15:41:03 +02:00 |
Miriam Baglioni
|
5feae77937
|
[Measures] last changes to accomodate tests
|
2022-04-20 15:13:09 +02:00 |
Miriam Baglioni
|
869407c6e2
|
[Measures] added new measure (usagecounts) as action set. Measure added at the level of the result. Ref #7587
|
2022-04-20 14:02:05 +02:00 |
Antonis Lempesis
|
b7cd2c6ca1
|
added open citations
|
2022-04-20 14:46:55 +03:00 |
Michele Artini
|
c96a8613f8
|
update SQL queries
|
2022-04-20 12:07:49 +02:00 |
Michele Artini
|
4314db55c8
|
migration to services: update sql queries
|
2022-04-19 15:05:02 +02:00 |
Miriam Baglioni
|
0012e57bf9
|
Merge branch 'master' of https://code-repo.d4science.org/D-Net/dnet-hadoop
|
2022-04-14 14:14:44 +02:00 |
Miriam Baglioni
|
c5a863132c
|
[BulkTagging] revert it
|
2022-04-14 14:14:13 +02:00 |
Sandro La Bruzzo
|
d5b29d96a7
|
fix merging in crossrefAggregator which creates dataInfo null
|
2022-04-14 11:07:04 +02:00 |
Miriam Baglioni
|
8e8933d41a
|
[BulkTagging] added fix if result.dataInfo is null
|
2022-04-14 09:04:24 +02:00 |
Claudio Atzori
|
b93a141d6c
|
[Doiboost] fixed fundingReference extraction from the Crossref records
|
2022-04-12 10:26:05 +02:00 |
Claudio Atzori
|
73c172926a
|
[Doiboost] fixed fundingReference extraction from the Crossref records
|
2022-04-12 10:25:42 +02:00 |
Claudio Atzori
|
48b580b45c
|
[graph enrichment] fixed country_propagation oozie workflow definition, parameter saveGraph is not needed anymore by the SparkCountryPropagationJob
|
2022-04-11 08:52:36 +02:00 |
Claudio Atzori
|
21f32b83c6
|
[graph enrichment] fixed country_propagation oozie workflow definition, parameter saveGraph is not needed anymore by the SparkCountryPropagationJob
|
2022-04-11 08:52:12 +02:00 |
Claudio Atzori
|
4eff7856f5
|
Merge pull request '[stats-wf] computing stats in each step' (#210) from antonis.lempesis/dnet-hadoop:beta into beta
Reviewed-on: D-Net/dnet-hadoop#210
|
2022-04-08 14:21:01 +02:00 |
Serafeim Chatzopoulos
|
d0b84d3297
|
Add FileCollectorPlugin and respective test
|
2022-04-07 15:06:38 +03:00 |
Serafeim Chatzopoulos
|
bc1bf55507
|
Add AbstractSplittedRecordPlugin
|
2022-04-07 14:33:04 +03:00 |
Claudio Atzori
|
c26222623f
|
[maven-release-plugin] prepare for next development iteration
|
2022-04-07 13:32:22 +02:00 |
Claudio Atzori
|
86585a6b27
|
[maven-release-plugin] prepare release dhp-1.2.4
|
2022-04-07 13:32:19 +02:00 |
Claudio Atzori
|
ad85d88eaf
|
[maven-release-plugin] rollback the release of dhp-1.2.4
|
2022-04-07 13:28:35 +02:00 |
Claudio Atzori
|
598e11dfd7
|
[maven-release-plugin] prepare for next development iteration
|
2022-04-07 13:27:02 +02:00 |
Claudio Atzori
|
db3d9877a5
|
[maven-release-plugin] prepare release dhp-1.2.4
|
2022-04-07 13:26:58 +02:00 |
Claudio Atzori
|
3bba6d6e38
|
[maven-release-plugin] rollback the release of dhp-1.2.4
|
2022-04-07 12:23:17 +02:00 |
Claudio Atzori
|
2ac2d928bd
|
[maven-release-plugin] prepare for next development iteration
|
2022-04-07 12:18:47 +02:00 |
Claudio Atzori
|
85bc722ff4
|
[maven-release-plugin] prepare release dhp-1.2.4
|
2022-04-07 12:18:43 +02:00 |
Claudio Atzori
|
bc05b6168a
|
[maven-release-plugin] rollback the release of dhp-1.2.4
|
2022-04-07 11:49:06 +02:00 |
Claudio Atzori
|
505420fd61
|
[maven-release-plugin] prepare for next development iteration
|
2022-04-07 11:34:06 +02:00 |
Claudio Atzori
|
66e718981e
|
[maven-release-plugin] prepare release dhp-1.2.4
|
2022-04-07 11:34:02 +02:00 |
Serafeim Chatzopoulos
|
e612489670
|
Add fileGZip collector plugin and respective test
|
2022-04-06 19:12:44 +03:00 |
Claudio Atzori
|
4190c9f6bc
|
[graph raw] avoid NPEs importing datasource consent fields
|
2022-04-06 15:34:31 +02:00 |
Claudio Atzori
|
05fafa1408
|
[graph raw] avoid NPEs importing datasource consent fields
|
2022-04-06 15:23:50 +02:00 |
Antonis Lempesis
|
c442c91f89
|
computing stats in each step
|
2022-04-06 12:40:02 +03:00 |
Claudio Atzori
|
8c457f1b2c
|
conflicts resolved, merged from beta
|
2022-04-06 10:27:52 +02:00 |
Miriam Baglioni
|
e77d104951
|
[OC] added / to workflow path
|
2022-04-05 15:07:11 +02:00 |
Miriam Baglioni
|
79336d46c5
|
[Clean Context] first naive implementation of a functionality to clean not wanted contextes from one result. This implementation simply verifies the main title of the results start with a given string
|
2022-04-04 15:52:31 +02:00 |
Claudio Atzori
|
873369af1c
|
Merge pull request '[stats wf] added apcs in monitor db' (#207) from antonis.lempesis/dnet-hadoop:beta into beta
Reviewed-on: D-Net/dnet-hadoop#207
|
2022-03-29 15:40:20 +02:00 |
Antonis Lempesis
|
7112806a73
|
views cannot be stored as parquet...
|
2022-03-29 16:37:29 +03:00 |
Antonis Lempesis
|
fff0b3cc19
|
added apcs in monitor db
|
2022-03-29 14:15:31 +03:00 |
Claudio Atzori
|
de85367695
|
Merge pull request '[stats wf] fix: views cannot be stored as parquet...' (#206) from antonis.lempesis/dnet-hadoop:beta into beta
Reviewed-on: D-Net/dnet-hadoop#206
|
2022-03-29 12:51:02 +02:00 |
Antonis Lempesis
|
ee24f3eb2c
|
views cannot be stored as parquet...
|
2022-03-29 13:47:48 +03:00 |
Sandro La Bruzzo
|
1b11010169
|
minor fix
|
2022-03-29 10:59:14 +02:00 |
Claudio Atzori
|
0a0ae84c22
|
[graph raw] DOI based instance URLs on https
|
2022-03-29 10:52:58 +02:00 |
Claudio Atzori
|
9fa3dd78fe
|
Merge pull request '[stats wf] various fixes, organization ids for inst. dashboard' (#205) from antonis.lempesis/dnet-hadoop:beta into beta
Reviewed-on: D-Net/dnet-hadoop#205
|
2022-03-28 22:03:49 +02:00 |
Claudio Atzori
|
96aa2a5d0d
|
Merge branch 'beta' into instance_group_by_url
|
2022-03-28 09:23:52 +02:00 |
Claudio Atzori
|
741bc99c47
|
Merge branch 'beta' into datasource_pdf_consent
|
2022-03-28 09:20:48 +02:00 |
Claudio Atzori
|
61319b2e83
|
updated dhp-schema version; set entity-level dataInfo before & after merging the fields from the group of duplicates
|
2022-03-25 16:38:33 +01:00 |
Antonis Lempesis
|
d8503cd191
|
added moooar organizations
|
2022-03-24 14:02:36 +02:00 |
Miriam Baglioni
|
7b8f85692e
|
[Enrichment country] fixed issues with parameters and workflow args
|
2022-03-23 17:20:23 +01:00 |
Claudio Atzori
|
48d32466e4
|
instances grouped by URL expose only one refereed
|
2022-03-23 14:52:03 +01:00 |
Claudio Atzori
|
f10066547b
|
increased spark.sql.shuffle.partitions in affiliation_from_semrel_propagation
|
2022-03-23 12:22:26 +01:00 |
Claudio Atzori
|
43733c1a18
|
Merge branch 'beta' of https://code-repo.d4science.org/D-Net/dnet-hadoop into beta
|
2022-03-23 12:14:27 +01:00 |
Antonis Lempesis
|
62f91b0869
|
cleanup
|
2022-03-22 16:17:49 +02:00 |
Antonis Lempesis
|
2e8394ecf8
|
creating aaall tables as parquet
|
2022-03-22 16:16:08 +02:00 |
Antonis Lempesis
|
dcfbeb8142
|
yet more typos
|
2022-03-21 12:36:03 +02:00 |
Miriam Baglioni
|
89fd275480
|
[HostedByMap] added left over from PR and fixed issue on workflow
|
2022-03-21 09:54:45 +01:00 |
miconis
|
c763aded70
|
dependency updated to the new pace-core version
|
2022-03-16 16:41:50 +01:00 |
miconis
|
c959639bd5
|
dependency updated to the new pace-core version
|
2022-03-15 16:33:03 +01:00 |
Miriam Baglioni
|
0f7d8ca2e0
|
[HostedByMap] change on master to align to PR 201 on beta merged as 9f3036c847
|
2022-03-11 15:16:02 +01:00 |
Claudio Atzori
|
f430029596
|
cleanup
|
2022-03-11 14:28:28 +01:00 |
Miriam Baglioni
|
12de9acb0d
|
[Country Propagation] left out from previous commit
|
2022-03-11 14:17:02 +01:00 |
Miriam Baglioni
|
2fbb35ade5
|
mergin with branch beta
|
2022-03-11 13:58:10 +01:00 |
Miriam Baglioni
|
4437f9345d
|
[Country Propagation] left out from previous commit
|
2022-03-11 13:57:47 +01:00 |
Miriam Baglioni
|
2b643059fa
|
[Country Propagation] changed the logic to get the collectedfrom at the result level. To fix issue when no instance is created for a result that should have the country associated. Change the code to use spark instead of hive to prepare the data needed for the propagation step. Added new tests for the intermediate steps and new verification for the propagation itself
|
2022-03-11 13:56:48 +01:00 |
Claudio Atzori
|
f25407bbe2
|
added mapping for datasource consent fields to integrate them in the graph
|
2022-03-11 09:32:42 +01:00 |
Miriam Baglioni
|
2c5087d55a
|
[HostedByMap] download of doaj from json, modification of test resources, deletion of class no more needed for the CSV download
|
2022-03-04 15:18:21 +01:00 |
Miriam Baglioni
|
5d608d6291
|
[HostedByMap] changed the model to include also oaStart date and review process that could be possibly used in the future
|
2022-03-04 11:06:09 +01:00 |
Miriam Baglioni
|
b7c2340952
|
[HostedByMap - DOIBoost] changed to use code moved to common since used also from hostedbymap now
|
2022-03-04 11:05:23 +01:00 |
Miriam Baglioni
|
8a41f63348
|
[HostedByMap] update to download the json instead of the csv
|
2022-03-04 10:38:43 +01:00 |
Miriam Baglioni
|
44b0c03080
|
[HostedByMap] update to download the json instead of the csv
|
2022-03-04 10:37:59 +01:00 |
Antonis Lempesis
|
ad78e505da
|
yet another fix
|
2022-03-03 12:28:12 +02:00 |
Miriam Baglioni
|
3be8737c32
|
[graph-stats] fixed query after the change in the indicator table related to PR#200
|
2022-03-02 14:09:05 +01:00 |
Miriam Baglioni
|
3970651ee1
|
Merge pull request 'fixed query after the change in the indicator table' (#200) from antonis.lempesis/dnet-hadoop:beta into beta
Reviewed-on: D-Net/dnet-hadoop#200
|
2022-03-02 14:05:58 +01:00 |
Antonis Lempesis
|
efeeebfee1
|
fixed query after the change in the indicator table
|
2022-03-02 13:29:25 +02:00 |
Claudio Atzori
|
580d904aae
|
manually merging PR#199 D-Net/dnet-hadoop#199
|
2022-02-25 12:22:50 +01:00 |
Claudio Atzori
|
1932a65d1c
|
Merge pull request '[Stats wf] sprint 6 indicators' (#198) from antonis.lempesis/dnet-hadoop:beta into beta
Reviewed-on: D-Net/dnet-hadoop#198
|
2022-02-25 12:09:18 +01:00 |
Miriam Baglioni
|
f5b0a6f89c
|
[master to beta] fixed issues in test files
|
2022-02-25 10:21:57 +01:00 |
miconis
|
8991d097b4
|
bug fix in the DedupRecordFactory, DataInfo set before merge
|
2022-02-24 17:13:12 +01:00 |
miconis
|
fe1c966cbf
|
Merge branch 'master_202203' of code-repo.d4science.org:D-Net/dnet-hadoop into master_202203
|
2022-02-24 17:08:38 +01:00 |
miconis
|
b0f369dc78
|
bug fix in the DedupRecordFactory, DataInfo set before merge
|
2022-02-24 17:08:24 +01:00 |
Miriam Baglioni
|
859cb7ac9d
|
[DoiBoost AR] changed test resource to be sure the result will always have EMBARGO as value for AccessRight
|
2022-02-24 16:55:32 +01:00 |
Miriam Baglioni
|
a40b59b7d5
|
[ResultToOrgFromInstRepoTest] fixed issue in model of the input resources
|
2022-02-24 16:05:57 +01:00 |
Claudio Atzori
|
66c09b1bc7
|
code formatting
|
2022-02-24 12:58:07 +01:00 |
Claudio Atzori
|
a87c070447
|
conflicts resolved, merged from beta
|
2022-02-24 12:51:31 +01:00 |
Claudio Atzori
|
86cdb7a38f
|
[provision] serialize measures defined on the result level
|
2022-02-23 15:54:18 +01:00 |
Alessia Bardi
|
9d6203f79b
|
test mapping datasource
|
2022-02-23 15:00:53 +01:00 |
Antonis Lempesis
|
3b92a2ab9c
|
added the rest of spring 6 in monitor db
|
2022-02-23 12:05:57 +02:00 |
Antonis Lempesis
|
87c91f70a2
|
added sprint 6 indicators to monitor db
|
2022-02-22 14:41:48 +02:00 |
Claudio Atzori
|
5226d0a100
|
Merge branch 'beta' of https://code-repo.d4science.org/D-Net/dnet-hadoop into beta
|
2022-02-18 15:21:07 +01:00 |
Claudio Atzori
|
99f5b14469
|
[graph raw] invisible records stored among the raw graph rather than the claimed subgraph
|
2022-02-18 15:20:57 +01:00 |
Claudio Atzori
|
401dd38074
|
code formatting
|
2022-02-18 15:19:15 +01:00 |
Claudio Atzori
|
cf8443780e
|
added processingchargeamount to the result view
|
2022-02-18 15:17:48 +01:00 |
Sandro La Bruzzo
|
891781ee3f
|
Merge branch 'beta' of code-repo.d4science.org:D-Net/dnet-hadoop into beta
|
2022-02-18 11:11:32 +01:00 |
Sandro La Bruzzo
|
d3f03abd51
|
fixed wrong json path
|
2022-02-18 11:11:17 +01:00 |
Claudio Atzori
|
89c7313fc5
|
Merge branch 'beta' into hierarchical_orgs_relations
|
2022-02-17 10:30:04 +01:00 |
dimitrispie
|
58c59f46eb
|
Added Sprint 6
|
2022-02-17 10:21:09 +02:00 |
Antonis Lempesis
|
5772f92dba
|
merged beta chnages in hive branch
|
2022-02-15 13:24:51 +02:00 |
Antonis Lempesis
|
393a4ee956
|
fixed yet another typo...
|
2022-02-15 12:56:50 +02:00 |
Sandro La Bruzzo
|
3aa2020b24
|
added script to regenerate hostedBy Map following instruction defined on ticket #7539
updated hosted By Map
|
2022-02-15 11:05:27 +01:00 |
Miriam Baglioni
|
be64055cfe
|
[OpenCitation] changed the name of destination folders
|
2022-02-14 15:49:44 +01:00 |
Miriam Baglioni
|
1490867cc7
|
[OpenCitation] cleaning of the COCI model
|
2022-02-14 14:52:12 +01:00 |
Miriam Baglioni
|
c191080965
|
mergin with branch beta
|
2022-02-14 14:49:39 +01:00 |
Alessia Bardi
|
600ede1798
|
serialisation of APCs int he XML records
|
2022-02-11 11:00:20 +01:00 |
Miriam Baglioni
|
5c4043dba8
|
[OpenCitation] refactoring
|
2022-02-08 16:23:05 +01:00 |
Miriam Baglioni
|
759ed519f2
|
[OpenCitation] added logic to avoid the genration of self citations relations
|
2022-02-08 16:15:34 +01:00 |
Miriam Baglioni
|
b071f8e415
|
[OpenCitation] change to extract in json format each folder just onece
|
2022-02-08 15:37:28 +01:00 |
Miriam Baglioni
|
fbc28ee8c3
|
[OpenCitation] change the integration logic to consider dois with commas inside
|
2022-02-07 18:32:08 +01:00 |
Miriam Baglioni
|
78be2975f0
|
[stats-wf]fixed another typo related to PR#193
|
2022-02-07 11:22:08 +01:00 |
Miriam Baglioni
|
1f8302dc37
|
Merge pull request '[stats-wf]fixed yet another typo' (#193) from antonis.lempesis/dnet-hadoop:beta into beta
Reviewed-on: D-Net/dnet-hadoop#193
|
2022-02-07 11:19:26 +01:00 |
Antonis Lempesis
|
5f762cbd09
|
fixed yet another typo
|
2022-02-07 12:09:12 +02:00 |
Alessia Bardi
|
ac8b8f224f
|
Merge branch 'beta' into extendResult
|
2022-02-04 16:43:27 +01:00 |
Miriam Baglioni
|
493caef358
|
[stats-wf]fixed the result_result table related to PR#191
|
2022-02-04 14:51:25 +01:00 |